Skip to content
Meta Low severity

Missing or Invalid HTML Lang Attribute

An indexable page whose root <html> tag has no lang attribute, or an invalid one, leaves browsers, screen readers and search engines guessing its language.

A screen reader guesses the wrong pronunciation, a browser offers to "translate this page" on a page already in the visitor's language, and neither has anything to do with your content. Both trace back to one missing attribute on the root tag.

What this finding means

Crawl Cove's html-lang check looks at every indexable page (HTTP 200, not noindex) that is an HTML document, and reads the lang attribute on the root <html> element. It flags the page in either of two cases: the attribute is missing or blank, or it is present but not a valid language tag.

A real finding looks like one of these:

Missing <html lang> attribute
Invalid lang value: "english"

Validity is checked against a pragmatic BCP-47 shape: a 2-3 letter primary subtag (en, de, pt) optionally followed by hyphen-separated subtags (en-US, pt-BR, zh-Hans). A value like english, a single letter, an underscore in place of a hyphen, or whitespace does not match and is flagged the same as an outright missing attribute.

Why it matters

The lang attribute tells browsers, screen readers and search engines what language the page is written in. Screen readers use it to select the correct pronunciation rules; get it wrong or leave it out, and a page in English can be read aloud with the phonetics of another language entirely. Browsers use it to decide whether to offer an automatic translation prompt, which fires unnecessarily, or fails to fire when it should, when the attribute is wrong or absent. Search engines lean on it too, alongside signals like hreflang, to understand which audience a page is written for. None of this is a direct ranking factor the way a broken canonical is, but it degrades the experience for exactly the visitors and assistive technology this attribute exists to serve.

How to fix it

  1. Open the flagged page's HTML source and find the root <html> tag.
  2. Add or correct the lang attribute: <html lang="en"> for English, or the appropriate BCP-47 code for your page's actual language (en-GB, fr, de, pt-BR).
  3. If your site serves multiple languages, make sure each template or CMS locale sets its own correct value rather than sharing one hardcoded default across every page.
  4. Re-crawl the page to confirm the finding clears.

False positives and edge cases

  • Only the root <html> element is checked. A lang attribute set on an individual element further down the page is a separate, narrower override this check does not evaluate either way.
  • Non-indexable pages are out of scope. A non-200 or noindex page is not being offered to search engines or most visitors, so it is not checked here.
  • Non-HTML resources are skipped. A PDF, image or other non-HTML response has no root <html> element to carry the attribute.
  • A technically valid but wrong language code still passes. This check confirms the value is a well-formed language tag, not that it correctly describes the page's actual content language.

Related reading

For the related international-SEO signal that tells search engines which language and region a URL targets, see the Hreflang Checker.

Frequently asked questions

What counts as a valid lang value?
A pragmatic BCP-47-style tag: a 2-3 letter primary subtag, optionally followed by one or more hyphen-separated subtags of 1-8 alphanumeric characters, for example en, en-US, pt-BR or zh-Hans. A value like english, e, en_US (underscore instead of hyphen) or a bare number fails validation and is flagged the same as a missing attribute.
Does the check look at more than the root tag?
No. It only reads the lang attribute on the root <html> element, the one declaration that applies to the whole document by default. A lang attribute set on an individual element further down the page is a different, narrower override this check does not evaluate.
Is a missing lang attribute a big ranking problem?
Not directly. Google does not use it the way it uses hreflang to pick which URL to show in a given region. Its real cost is to screen readers, which rely on it to choose the correct pronunciation, and to browsers, which use it to decide whether to offer an automatic translation prompt.
Does this run on non-HTML pages?
No. A PDF, image or other non-HTML resource has no root <html> element to carry the attribute, so it is out of scope.

Audit your site the easy way

Crawl Cove finds this on your machine, on every plan, and tells you exactly what to fix first. See the features or compare the plans.

Download Crawl Cove