robots.txt

robots.txt is a file at the root of a site that tells automated crawlers which paths they may request, using per-agent rules.

It controls crawling, not indexing. A page disallowed in robots.txt can still appear in search results if other pages link to it, because the crawler is told not to fetch it rather than not to list it. To keep a page out of an index, a noindex directive on the page itself is the correct mechanism, and that only works if the crawler is allowed to fetch the page and see it.

The most expensive robots.txt mistake is a blanket rule that catches an AI retrieval agent nobody meant to block. The second most expensive is a file that declares nothing at all, which is what a default managed file usually amounts to.

Check your robots.txt against every AI agent

Related

Tell us about your business.

Thirty minutes with a specialist. You leave with a plan, whether or not you hire us.