[Atomic Glue](atomicglue.co)
SEO
Home › Glossary › SEO· 352 ·

Robots.txt

\ro-bots dot tee-eks-tee\n.
Filed underSEO
In brief · quick answer

Robots.txt is a text file placed at the root of a website that instructs search engine crawlers which URLs they may or may not access.

§ 1 Definition

Robots.txt is the first thing Googlebot checks when it arrives at a site. It is a plain text file at `https://example.com/robots.txt` that uses a simple syntax to allow or disallow crawling of specific paths. It is a crawl control mechanism, not a security tool. Blocked URLs are still accessible to anyone who knows the URL; robots.txt only stops bots that choose to obey it.

§ 2 Common Robots.txt Directives

Use `User-agent` to specify which bot the rule targets (`*` means all bots). Use `Disallow` to block paths. Use `Allow` to override a broader Disallow (for specific files within a blocked directory). Use `Sitemap` to point to your XML sitemap location. Use `Crawl-delay` to suggest a crawl rate (though Googlebot ignores this in favor of actual server response times).

§ 3 The Robots.txt Trap

The most dangerous robots.txt mistake is accidentally blocking Googlebot from crawling your entire site. A `Disallow: /` rule for all user-agents means Googlebot cannot crawl anything. A `Disallow:` rule with no path means nothing is blocked. This distinction has tripped up many site migrations.

§ 4 Note

Robots.txt blocks crawling but does not prevent indexing. If another site links to a blocked URL, Google may still index it (just without crawling the content). For actual index blocking, use a `noindex` meta tag or X-Robots-Tag header. This is one of the most misunderstood distinctions in technical SEO.

§ 5 In code

User-agent: *
Disallow: /wp-admin/
Disallow: /private/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/sitemap.xml
Key takeaways
  • Robots.txt controls crawling, not indexing
  • A single misconfigured rule can block entire site indexing
  • Use both robots.txt (crawl control) and meta robots (index control)
How Atomic Glue helps

Atomic Glue builds SEO strategies that turn technical fundamentals into measurable rankings. Get in touch to discuss your SEO & GEO services needs.

Get in touch

Robots.txt is a text file placed at the root of a website that instructs search engine crawlers which URLs they may or may not access.

Definition

Robots.txt is the first thing Googlebot checks when it arrives at a site. It is a plain text file at https://example.com/robots.txt that uses a simple syntax to allow or disallow crawling of specific paths. It is a crawl control mechanism, not a security tool. Blocked URLs are still accessible to anyone who knows the URL; robots.txt only stops bots that choose to obey it.

Common Robots.txt Directives

Use User-agent to specify which bot the rule targets (* means all bots). Use Disallow to block paths. Use Allow to override a broader Disallow (for specific files within a blocked directory). Use Sitemap to point to your XML sitemap location. Use Crawl-delay to suggest a crawl rate (though Googlebot ignores this in favor of actual server response times).

The Robots.txt Trap

The most dangerous robots.txt mistake is accidentally blocking Googlebot from crawling your entire site. A Disallow: / rule for all user-agents means Googlebot cannot crawl anything. A Disallow: rule with no path means nothing is blocked. This distinction has tripped up many site migrations.

Note

Robots.txt blocks crawling but does not prevent indexing. If another site links to a blocked URL, Google may still index it (just without crawling the content). For actual index blocking, use a noindex meta tag or X-Robots-Tag header. This is one of the most misunderstood distinctions in technical SEO.

In code


User-agent: *
Disallow: /wp-admin/
Disallow: /private/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://example.com/sitemap.xml

Key takeaways:

  • Robots.txt controls crawling, not indexing

  • A single misconfigured rule can block entire site indexing

  • Use both robots.txt (crawl control) and meta robots (index control)

Schedule a call

30 min · Video call

1
Date
2
Time
3
Details