Node CLI · v1.0.0 · MIT
A schema validator for the terminal and CI
crawlcove-schema-validator is the command line companion to the schema markup generator on this site: the generator writes the JSON-LD, this checks what you published. Give it a URL, a local HTML file, or a bare JSON-LD document on stdin, and it extracts every application/ld+json block, parses it and names the character at fault when it does not parse, flattens @graph, and checks each node against the properties Google's rich-result documentation requires and recommends for Article, Product, Offer, FAQPage, BreadcrumbList, Organization, LocalBusiness, Event, Recipe, HowTo, JobPosting, VideoObject, SoftwareApplication and more. It also catches the mistakes that are provably mistakes: a lower-case type, a missing @context, a relative URL, a date that is not ISO 8601, a breadcrumb with gaps in its positions. Exit code 1 in CI on an error.
Free · MIT-licensed · README checked September 2026
Install and run
# one-off, nothing installed (Node 18+):
npx github:CrawlCove/crawlcove-schema-validator https://www.example.com/product/widget
# global command, from the release tarball:
npm install -g https://github.com/CrawlCove/crawlcove-schema-validator/archive/refs/tags/v1.0.0.tar.gz
schema-validator --version
Options
schema-validator <url> [options]
schema-validator --file page.html
cat markup.json | schema-validator --file -
-f, --file <path> validate a local HTML file or a bare JSON-LD document ("-" for stdin)
--timeout <ms> request timeout (default 10000)
--user-agent <ua> User-Agent header to send
--json JSON output
--fail-on <level> error (default), warning, none
The README in the repo is the authoritative reference; the commands above are copied from it as of September 2026. If they disagree, the README is newer.
What it checks
- invalid-json: the block does not parse, so search engines ignore all of it. The message is the parser's own reason; the usual cause is a raw line break inside a string.
- missing-context, missing-type, type-case and malformed-type: no @context, no @type, product instead of Product, or a URL where a bare type name belongs.
- missing-required, as an error: a property Google lists as required for that type's rich result. Subtypes such as BlogPosting or Restaurant are checked against their parent's rules.
- missing-recommended, as a warning, raised only for top-level nodes, so a nested author Person is not nagged about its jobTitle.
- wrong-nested-type, breadcrumb-positions, invalid-date, invalid-duration and relative-url: the structural mistakes that quietly disqualify a page from a rich result.
- duplicate-type, empty-value and empty-block, plus no-structured-data and microdata-only as warnings when there is nothing to validate.
Limits worth knowing
- It validates JSON-LD only. Microdata is reported as present but not checked; Google recommends JSON-LD anyway.
- Types it has never heard of are left alone. Schema.org has about 800, and a rare one is more likely real than wrong, so an unknown type is never an error.
- A clean result means the markup meets Google's documented property rules, not that the page will show a rich result; eligibility also depends on content and on Google's own judgement.
Works with Crawl Cove
Schema validator CLI covers the checks it lists above. The desktop app runs the full site audit: every page, every finding ranked by impact, history over time and Search Console data alongside. See the desktop checker or compare the plans.
Frequently asked questions
- How does this differ from Google's Rich Results Test?
- The Rich Results Test is a web page you paste one URL into and read by eye. This is a command you can run on a hundred URLs from a script, on a file before it is deployed, or on stdin from another tool, and it returns an exit code. Use Google's test for the final word on a live page; use this to stop the broken block reaching the live page.
- What does it check against?
- The required and recommended properties from Google's rich-result documentation for more than thirty types, with subtypes mapped to their parents. Required properties are errors; recommended ones are warnings, raised only for top-level nodes.
- Can I validate markup before it is on a page?
- Yes. Pass a local HTML file with --file, or pipe a bare JSON-LD document to --file -, which is how you check the output of the schema markup generator on this site before pasting it into a template.
- Does it check microdata or RDFa?
- No. It reports microdata as present so you know it is there, and warns when a page has only microdata, but validates JSON-LD blocks only. Google recommends JSON-LD, and moving to it is the fix the warning suggests.
- Can it check every page on a site?
- It validates the pages you name, one run each. To find every page on a site that should carry structured data and does not, or whose markup is broken, crawl the site with Crawl Cove, which runs the same schema-validity check across the whole crawl and tracks it over time.
Stop guessing.
Start fixing.
Crawl Cove runs on your machine, connects to your real ranking data, and tells you exactly what to fix first. No per-feature paywalls, no spreadsheets, no guesswork.
28 days risk-free · No card required