Guides

Test robots.txt rules with real paths

Robots.txt controls crawling for its own protocol, host and port. Place it at the root, such as https://example.com/robots.txt. A rule in a different folder is not the root crawl policy. Crawlers may cache an earlier version, so a local test can differ from an actual crawl.

Choose a crawler token such as Googlebot and enter a case-sensitive path. More specific user-agent groups take precedence over the wildcard group. Groups for the same selected agent are combined. Within those groups, the most specific matching rule wins; allow wins an equal-specificity tie.

For example, Disallow: /private/ blocks that directory, while Allow: /private/public$ permits the exact public path. The * wildcard matches any sequence and $ anchors the end. The query string is part of the tested path. Empty disallow rules impose no restriction.

Revbu shows matching rules and their original lines, plus unsupported fields. Google does not support crawl-delay in this format. Review the actual bot documentation if you test another crawler. This simulator does not predict caching, server errors or crawler-specific extensions.