robots.txt: The Complete Guide
robots.txt tells crawlers where they may and may not go. One wrong line can remove your whole site from search — here is how to get it right.
What robots.txt is for
robots.txt is a plain-text file at your domain root that tells crawlers which parts of your site they may fetch. It is a set of requests, not a security control: reputable bots obey it, malicious ones ignore it.
It is not for hiding pages from search results — use noindex for that.
The essential directives
| Directive | Meaning |
|---|---|
User-agent |
Which crawler the block applies to (* = all) |
Allow |
Paths that may be crawled |
Disallow |
Paths that may not be crawled |
Sitemap |
Full URL of your XML sitemap |
Crawl-delay |
Seconds between requests (often ignored) |
A safe starter file
User-agent: *
Allow: /
Disallow: /api/
Disallow: /search
Disallow: /*?s=
Sitemap: https://example.com/sitemap.xml
Generate a correct file in seconds with the robots.txt generator.
Allow vs Disallow order
Longest-match wins in modern crawlers, so specificity matters. Allow: /blog can carve out an exception under Disallow: /. Test ambiguous rules before deploying.
Six mistakes that hurt sites
Disallow: /on production — this deindexes your whole site.- Blocking CSS and JS — Google needs them to render pages and judge mobile-friendliness.
- Relying on robots.txt for privacy — use authentication instead.
- Blocking the sitemap itself or the crawl of pages you want ranked.
- Typos in directives —
Disallow /(no colon) is silently ignored. - No sitemap reference — a missed opportunity to guide crawling.
robots.txt and noindex together
This is the classic SEO trap: if a page is disallowed in robots.txt, the crawler never sees its noindex tag — so the page can still be indexed. To remove a page from the index, allow crawling and add noindex, or use the URL removal tool.
Test before you ship
Deploy to staging, fetch /robots.txt, and confirm expected URLs are allowed. Then check coverage in Search Console after release.
Build yours
Use the free robots.txt generator, then pair it with an XML sitemap so crawlers find every important page.