How Google matches rules
- The crawler uses the group whose
User-agentmatches its name most specifically;Googlebot-Imagefalls back to aGooglebotgroup, then to*. Groups for the same agent are merged. - Within that group, the rule with the longest path that matches wins.
*matches any run of characters and$anchors the end. - If an
Allowand aDisalloware equally long,Allowwins.
What robots.txt does not do
It stops crawling, not indexing. A blocked URL can still appear in results if other pages link to it — without a description, because Google could not read the page. To keep a page out of the index, let it be crawled and use noindex. The tester flags directives Google ignores, such as crawl-delay and noindex inside robots.txt.
AI crawlers
GPTBot, ClaudeBot, PerplexityBot and Google-Extended each read their own group. Test them by name to see whether your file lets AI products train on or quote your pages.