Technical SEO

robots.txt: The Complete Guide

robots.txt tells crawlers where they may and may not go. One wrong line can remove your whole site from search — here is how to get it right.

Updated 16 September 2026 · 8 min read
Advertisement

What robots.txt is for

robots.txt is a plain-text file at your domain root that tells crawlers which parts of your site they may fetch. It is a set of requests, not a security control: reputable bots obey it, malicious ones ignore it.

It is not for hiding pages from search results — use noindex for that.

The essential directives

Directive Meaning
User-agent Which crawler the block applies to (* = all)
Allow Paths that may be crawled
Disallow Paths that may not be crawled
Sitemap Full URL of your XML sitemap
Crawl-delay Seconds between requests (often ignored)

A safe starter file

User-agent: *
Allow: /
Disallow: /api/
Disallow: /search
Disallow: /*?s=

Sitemap: https://example.com/sitemap.xml

Generate a correct file in seconds with the robots.txt generator.

Allow vs Disallow order

Longest-match wins in modern crawlers, so specificity matters. Allow: /blog can carve out an exception under Disallow: /. Test ambiguous rules before deploying.

Six mistakes that hurt sites

  1. Disallow: / on production — this deindexes your whole site.
  2. Blocking CSS and JS — Google needs them to render pages and judge mobile-friendliness.
  3. Relying on robots.txt for privacy — use authentication instead.
  4. Blocking the sitemap itself or the crawl of pages you want ranked.
  5. Typos in directives — Disallow / (no colon) is silently ignored.
  6. No sitemap reference — a missed opportunity to guide crawling.

robots.txt and noindex together

This is the classic SEO trap: if a page is disallowed in robots.txt, the crawler never sees its noindex tag — so the page can still be indexed. To remove a page from the index, allow crawling and add noindex, or use the URL removal tool.

Test before you ship

Deploy to staging, fetch /robots.txt, and confirm expected URLs are allowed. Then check coverage in Search Console after release.

Build yours

Use the free robots.txt generator, then pair it with an XML sitemap so crawlers find every important page.

Advertisement
Help Center

robots.txt: The Complete Guide — FAQ

At the root of your domain: https://example.com/robots.txt. It cannot live in a subfolder, and it must be publicly reachable.

No. It prevents crawling, not indexing. A disallowed URL can still appear in results if it is linked elsewhere. To keep a page out of the index, use a noindex tag or meta robots instead.

It tells well-behaved crawlers not to fetch anything on the site. Use it only on staging or private sites — it will remove your site from search.

Keep reading

cta.title

cta.text

cta.button
100% Free No signup Private & secure
Advertisement