Skip to content
Developer tools 7 min read By The Crawl Cove team

SEO Crawl Export Format: JSON Schema and CSV

What is inside a Crawl Cove crawl export, field by field, how JSON and CSV differ, and how to validate one against the open-source JSON Schema.

Key takeaways

  • A crawl export is only useful to a script if its format is written down, and crawlcove-export-spec is that document: a JSON Schema, a CSV column reference, a versioning policy and a real sample file
  • The JSON export is six top-level fields and a pages array; each page carries 29 required fields, from URL and status to content fingerprint and findings count
  • The CSV carries the same pages and fields in a fixed column order, but booleans become yes/no and every null becomes an empty cell, so the JSON is the source of truth where that difference matters
  • Validate with the bundled Ajv script or compile the schema in your own code, and pin the schemaVersion your tooling was written against
  • The CLI and the GitHub Action use the same field names where they check the same thing, so one reader covers all three

Every crawler can export. Very few tell you what the export means precisely enough to write code against it, which is why so many "import your crawl" scripts start with an afternoon of reverse-engineering column headers. Since version 1.2.0, the Crawl Cove desktop app exports a report's full dataset from the Reports page as JSON or CSV, and crawlcove-export-spec is the document that makes that export something you can build on: a JSON Schema, a CSV column reference, a versioning policy and a real sample file. This post walks through the format.

The shape of the JSON export

The JSON export is one object with six top-level fields, all required:

Field What it holds
schemaVersion Semver of the schema the file follows, such as 1.0.0. Bumped on any added, renamed or retyped field.
exportedAt When the export was produced.
auditRunId The identifier of the crawl run inside the app, so two exports of the same run can be recognised as such.
runIncomplete Whether the crawl was stopped before it finished. A true value means the pages array is a partial picture.
pageCount The number of entries in pages.
pages The array of page records.

The first thing a reader should do is check schemaVersion against the version it was written for, and the second is check runIncomplete, because a partial crawl will make every site-wide count look better than it is.

The page record, field by field

Each entry in pages carries 29 required fields. Grouped by what they describe:

Identity and response. url is the URL as requested; finalUrl is where it ended up after redirects, or null if not recorded. statusCode is the final response's HTTP status, null if the request never got one. contentType, depth (0 for the start URL), responseTimeMs, byteSize and fetchError complete the picture of the request itself. redirectHops is the number of redirects the URL went through, so a value of 2 or more is a redirect chain.

Indexability. indexable is true only when the page returned 200 and is not noindex by either the robots meta tag or the X-Robots-Tag header. Both raw values are kept alongside it, as robotsMeta and xRobotsTag, so a script can see which one made the decision. canonical is the declared canonical href, null if absent.

On-page. title and titleLength, metaDescription and metaLength, h1Count, wordCount, htmlLang. The length fields are character counts and are null when the tag is absent, which is not the same as zero.

Links and media. linksInternal and linksExternal are counts. imageCount and imagesMissingAlt are what an image alt text audit starts from.

Structure. schemaBlocks counts the JSON-LD blocks on the page, hreflangCount the rel="alternate" hreflang entries, and rendered records whether the page was fetched with a JavaScript rendering pass rather than the plain HTTP response.

Crawl Cove's own. contentFingerprint is a 64-bit fingerprint of the page's content as hex, used for near-duplicate detection; it is an empty string, not null, when the page had too little content to fingerprint. findingsCount is how many of the run's SEO findings were attributed to this page.

That last distinction, empty string versus null, is one the schema is careful about, and it is the reason the JSON export is the source of truth rather than the CSV.

What changes in the CSV

The CSV export has the same page set and the same fields in a fixed column order, from URL in column 1 to Findings in column 29. It is RFC 4180: comma-separated, CRLF line endings, UTF-8 with a byte-order mark so Excel opens it with the right encoding. Three things are re-encoded on the way out, because CSV has no native booleans or nulls:

  • Indexable and Rendered become the strings yes and no.
  • Every null becomes an empty cell. In a CSV cell there is no way to tell "empty string" from "absent", so the contentFingerprint case above collapses in the CSV.
  • Any cell whose content would parse as a spreadsheet formula, meaning it starts with =, +, -, @, a tab or a carriage return, is prefixed with a single quote so Excel, Sheets and LibreOffice treat it as text rather than executing it. That can only affect the free-text fields sourced from the crawled page: Title, Meta description, Fetch error, URL, Final URL and Canonical.

Note

If you are building anything more than a one-off spreadsheet, read the JSON. The CSV is the right export for a human with a spreadsheet; the JSON is the right one for code, and the schema describes the JSON.

Validating an export

The repo ships a small Ajv-based validator. Clone it, install, and point it at your file:

git clone https://github.com/CrawlCove/crawlcove-export-spec.git
cd crawlcove-export-spec
npm install
npm run validate -- path/to/your-export.json

Run it without a path and it validates the bundled sample, which is worth doing once: the sample is a real export produced by the app's own export code against a small seeded crawl, not hand-written JSON, so it is a true instance of the format.

In your own code, compile the schema and call the validator on the parsed export:

const Ajv = require('ajv')
const addFormats = require('ajv-formats')
const schema = require('./schema/crawl-export.schema.json')

const ajv = new Ajv()
addFormats(ajv)
const validate = ajv.compile(schema)

if (!validate(myExport)) {
  console.error(validate.errors)
}

validate.errors names the path and the rule for every mismatch, which is the fastest way to find out that a field you assumed was a number is sometimes null.

Versioning, and why to pin it

schemaVersion is bumped on any added, renamed or retyped field. The versioning document in the repo says which of those bumps break an existing reader. The practical rule is the same as for any data contract: record the version your tooling was written against, check the field on every file you read, and read the versioning document before you let a newer version through.

One reader for three tools

The export spec is the shared vocabulary for the rest of Crawl Cove's open-source tools. The crawlcove-cli crawler's per-page output uses this spec's field names wherever it checks the same thing the app does; the GitHub Action writes its report in the same shape; and the MCP server can load a desktop app export from disk with its load_export tool and answer questions about it without the live crawl's 200-page limit. A script that reads one of these reads all of them. The format is also the landing point for a crawl made elsewhere: crawlcove-sf-import converts a Screaming Frog Internal export into it, fills the 22 of its 29 page fields that the Frog's columns cover, leaves the rest null with a report saying why, and validates the result against this spec, so a Frog crawl can go through the same reader as everything above. Two of those readers are ready-made: crawlcove-js is a typed, zero-dependency JavaScript library that loads either format and answers questions like which titles are duplicated, and crawlcove-sheets imports the same export into a Google Sheets audit workbook. Both apply the same issue rules as the MCP server, so a script, a sheet and an assistant built on one export agree.

To produce an export yourself, run a crawl in the Crawl Cove desktop app and use the export on the Reports page; exporting your data covers where the buttons are. For what to do with the findings once you have them, the technical SEO audit checklist is the working order.

Frequently asked questions

What is the SEO crawl export format?
The documented shape of the file the Crawl Cove desktop app produces when you export a report's full dataset from the Reports page. crawlcove-export-spec publishes it as a JSON Schema (2020-12) for the JSON export and a column reference for the CSV export, with a sample file and a validator.
Which version of Crawl Cove produces it?
Crawl Cove 1.2.0 and later. Earlier versions exported findings and keywords as CSV but not the full per-page dataset.
What is in each page record?
29 required fields: url, finalUrl, statusCode, contentType, depth, indexable, title, titleLength, metaDescription, metaLength, canonical, htmlLang, robotsMeta, xRobotsTag, h1Count, wordCount, linksInternal, linksExternal, imageCount, imagesMissingAlt, responseTimeMs, byteSize, rendered, redirectHops, fetchError, schemaBlocks, hreflangCount, contentFingerprint and findingsCount.
How do JSON and CSV differ?
The CSV has the same pages and fields in a fixed column order, but CSV has no booleans or nulls. indexable and rendered become the strings yes and no, and every null becomes an empty cell, which cannot be told apart from an empty string. Cells that would parse as a spreadsheet formula are prefixed so they are treated as text.
How do I validate an export?
Clone the repo, run npm install, then npm run validate with the path to your file. Without a path it validates the bundled sample. In your own code, compile schema/crawl-export.schema.json with Ajv and ajv-formats and call the validator on the parsed export.
What counts as a breaking change?
schemaVersion is bumped on any added, renamed or retyped field. The versioning document in the repo says which bumps break existing readers, so pin the version your tooling expects and read it before upgrading.

Audit your site the easy way

Crawl Cove finds these issues on your machine and tells you exactly what to fix first. See the features or compare the plans.

Download Crawl Cove