robots.txt
robots.txt is a file at the root of a site that tells automated crawlers which paths they may request, using per-agent rules.
It controls crawling, not indexing. A page disallowed in robots.txt can still appear in search results if other pages link to it, because the crawler is told not to fetch it rather than not to list it. To keep a page out of an index, a noindex directive on the page itself is the correct mechanism, and that only works if the crawler is allowed to fetch the page and see it.
The most expensive robots.txt mistake is a blanket rule that catches an AI retrieval agent nobody meant to block. The second most expensive is a file that declares nothing at all, which is what a default managed file usually amounts to.
Related
AI crawler
An AI crawler is an automated fetcher operated by a model provider that reads web pages either to answer a user's question in real time or to build training data.
Canonical tag
A canonical tag is a link element that names the preferred URL for a page, telling search engines which version to index when the same content is reachable at more than one address.
Tell us about your business.
Thirty minutes with a specialist. You leave with a plan, whether or not you hire us.