Key takeaways
- A robots.txt file controls whether a crawler may fetch a URL, and nothing about whether that URL can appear in Google's results
- Disallowing a page does not remove it from Google; a disallowed page can still be indexed if something else links to it, and blocking a page you also want deindexed stops Google reading the noindex tag that would actually remove it
- Crawlers obey the most specific user-agent group that matches them, then the longest matching path inside it, and Allow wins an exact-length tie with Disallow
- Google retired its standalone robots.txt Tester from Search Console at the end of 2023; the current way to check a live file is the robots.txt report under Indexing
A robots.txt file is a handful of lines of plain text, and most of the ones we've seen live on real sites get at least one line wrong: a Disallow: / left over from staging, a rule that silently blocks the CSS Google needs to render the page, or a block aimed at hiding a page from search that does the opposite of what was intended. Here's how the syntax actually works, how a crawler picks between competing rules, the mistakes that cost the most, and how to test a file before and after you publish it.
What a robots.txt file actually does
Google's own documentation is direct about the scope of this file:
"A robots.txt file tells search engine crawlers which URLs the crawler can access on your site. This is used mainly to avoid overloading your site with requests."
That's it. Robots.txt is a crawling control, not an indexing control, and the difference between those two words is where most of the expensive mistakes on this file come from. It tells a well-behaved crawler which URLs it may fetch. It says nothing about whether a URL is allowed to appear in search results, and it has no effect at all on crawlers that choose to ignore it.
Basic syntax
A robots.txt file is organised into groups, each starting with one or more User-agent lines followed by the rules that apply to them:
User-agent: Googlebot
Disallow: /admin/
Allow: /admin/help/
User-agent: *
Disallow: /admin/
Disallow: /search
Sitemap: https://example.com/sitemap.xml
User-agentnames which crawler the group applies to, by the token it identifies itself with (Googlebot,Bingbot,GPTBot, and so on), or*as a fallback that matches any crawler with no more specific group of its own.Disallownames a path prefix the crawler should not fetch.Disallow: /admin/blocks that folder and everything under it.Allowcarves out an exception inside a broaderDisallow, as in the example above, where/admin/help/stays fetchable even though the rest of/admin/is blocked.Sitemappoints to your sitemap's full URL. It can sit in any group, or on its own, and it's the one directive every crawler reads regardless of whichUser-agentgroup it belongs to.- Wildcards work in paths:
*matches any sequence of characters, and$anchors the match to the end of the URL, soDisallow: /*.pdf$blocks every PDF without also blocking a URL that merely contains.pdfpartway through.
Rules are case-sensitive, and the file must be plain text in UTF-8. A file produced by a word processor rather than a plain-text editor can carry invisible curly quotes or other characters that silently break a rule; if a rule you wrote isn't behaving as expected, checking for exactly this is a cheap first step.
How a crawler picks between rules
A site of any size ends up with more than one group and more than one matching rule, and crawlers resolve conflicts in a fixed order rather than reading top to bottom:
- Most specific user-agent group wins. A crawler that matches a named group (
User-agent: Googlebot) uses that group's rules and ignores the wildcard*group entirely, even if the wildcard group would otherwise have blocked the same path. - Inside that group, the longest matching path wins. If
Disallow: /blog/andAllow: /blog/2026/both match a URL under/blog/2026/, the longer, more specific/blog/2026/rule decides the outcome. - On an exact-length tie, Allow wins. If an
Allowand aDisallowrule match the same path with equal specificity, the crawler treats the URL as allowed.
Getting this wrong is easy on a file with more than a handful of rules, because the outcome for any one URL depends on which rule is most specific, not which rule is written first or last. Our free robots.txt generator applies exactly this resolution order and tells you which line decided the result for any URL and crawler you test, rather than leaving you to trace it by hand.
The mistake that costs the most: Disallow is not noindex
This is the single most expensive misunderstanding about this file, and Google states the correction plainly:
"It is not a mechanism for keeping a web page out of Google. To keep a web page out of Google, block indexing with
noindexor password-protect the page."
Google's documentation adds the specific failure mode: "A page that's disallowed in robots.txt can still be indexed if linked to from other sites." Google won't fetch a blocked page's content, but if another site links to it, the URL and the linking site's own anchor text can still surface in results, frequently as a bare listing with no description, which reads as more alarming than a normal search result but is exactly what Google's documentation describes as expected behaviour, not a bug.
The two mechanisms are mutually exclusive on the same URL, and the reason is mechanical rather than a policy choice: a noindex tag has to be read by a crawler that fetches the page. Disallow it in robots.txt at the same time, and the crawler that would have read the noindex tag never gets there, so the block you wanted never happens. If a page needs to come out of Google entirely, the sequence matters: allow the crawl, serve noindex, wait for Google to recrawl and drop it, and only add a Disallow afterwards if you also want to stop the crawl traffic once the page is genuinely gone.
Other mistakes worth checking for
- A leftover
Disallow: /from a staging environment. This single line blocks every URL on the site, homepage included, for whicheverUser-agentgroup it sits in. It's the most common way a site accidentally deindexes itself after a migration or a redesign, and it's worth checking for explicitly after every deploy, not just assuming it isn't there. - Blocking CSS or JavaScript. Googlebot renders pages the way a browser does before evaluating them, so blocking the resources a page needs to render can mean Google sees a broken or blank layout instead of the finished page, which can cost rankings for reasons that have nothing to do with your actual content.
- Wrong location. The file only works at the root of the host it needs to cover:
example.com/robots.txt, not a subfolder. A subdomain needs its own file at its own root; a file at the main domain's root has no effect on a subdomain's crawling at all. - Forgetting the
Sitemapline. It costs one line, is read regardless of whichUser-agentgroup it's in, and gives every well-behaved crawler a direct pointer to your full URL list instead of relying on discovery through links alone. - Assuming every crawler obeys it. Robots.txt is a voluntary standard. Well-behaved crawlers, including Google's, respect it. A crawler that ignores it entirely will not be stopped by anything written in this file.
How to test it
Search Console's standalone robots.txt Tester, the tool most guides still describe, was retired at the end of 2023. For a file already live on your site, Search Console's robots.txt report, under Indexing, is the current equivalent. For testing rules before you publish, or for a URL-by-URL check against a specific crawler without waiting on Search Console, Google also publishes its own robots.txt parser as open source, which is the exact logic Google's crawlers apply.
Our free robots.txt generator covers both sides of this in one place: fetch your live file and test any URL against Googlebot, Bingbot, GPTBot, and more to see whether it's allowed or blocked and which line decided it, or build a new one from scratch. If you're starting a file for a common platform rather than from a blank page, it includes real starter templates for WordPress, Joomla, Magento and Drupal, plus one-click presets for the rules people reach for most (adding a sitemap line, blocking AI bots as a group, or allowing only Googlebot).
If AI crawlers specifically are what you're trying to control, allow or block, that's a big enough topic on its own, covered in AI crawlers in robots.txt: who to allow and block, including the actual user-agent tokens for GPTBot, ClaudeBot, PerplexityBot and the rest.
Seeing it across a whole site, not one file at a time
Writing and testing the file itself only answers half the question. The other half is whether the rules you wrote actually match what you intended once they're live: is an important section accidentally caught by a broad Disallow, is a page you meant to keep crawlable actually blocked, is the sitemap line even being read. Crawl Cove crawls your whole site and flags every important page your robots.txt or a stray noindex tag is quietly keeping out, rather than you checking pages one URL at a time.
Wrap-up
A robots.txt file is a small amount of syntax with an outsized ability to cause damage when it's wrong, mostly because its one job, controlling crawl access, gets confused with a job it was never built for, controlling what appears in search results. Get the syntax right, understand that the most specific rule wins rather than the first or last one written, keep Disallow and noindex doing their separate jobs rather than fighting each other on the same URL, and test the result against a real crawler before you trust it.
Frequently asked questions
- Where does robots.txt need to live?
- At the root of the domain only, for example crawlcove.com/robots.txt. A file at crawlcove.com/blog/robots.txt or on a subdomain it wasn't published to is invisible to crawlers checking that host; each subdomain needs its own file at its own root if it needs different rules.
- Does Disallow stop a page ranking in Google?
- Not reliably. Disallow only stops Google fetching the page. If another site links to a disallowed URL, Google can still show that URL in results using the link text and surrounding context, without ever having crawled the page itself, often as a bare listing with no description.
- What actually removes a page from Google's results?
- A noindex directive, either as a meta robots tag or an X-Robots-Tag HTTP header, or a password wall. Both require Google to fetch the page to read the instruction, which is exactly what a Disallow rule prevents, so the two are mutually exclusive on the same URL: block the crawl or serve noindex, never both on a page you want gone.
- Can I use robots.txt to keep a page out of Google entirely?
- No. Google's own documentation states plainly that robots.txt "is not a mechanism for keeping a web page out of Google." Use noindex or a password wall instead.
- How do I test a live robots.txt file today?
- Search Console's standalone robots.txt Tester was retired at the end of 2023. The current options are the robots.txt report under Indexing in Search Console for a file already live, or Google's own open-source robots.txt parser library for testing rules before you publish them.