Robots.txt AI crawler checker

We fetch the public robots.txt file. The result applies to the selected path, which defaults to the homepage.

Use this robots.txt AI crawler checker to inspect published rules for each listed crawler or control token.

Read the rule for the page you care about

The checker requests the robots.txt file at the root of the domain you enter. It evaluates the selected path, which starts as the homepage. Change that path when you want to inspect a product, pricing or help page. A site can allow its homepage while disallowing a deeper directory, so the path is part of the result.

The table shows allow, block or not specified for each token, together with the matching group and deciding line where one exists. Open the checked file to review its context. Use that evidence when discussing a change with the person who maintains your website or crawler policy.

Understand the three results

Allow means a matching Allow directive wins for the selected path. Block means a matching Disallow directive wins. Not specified means no directive matches that path in the applicable group; ordinary robots processing permits it by default. The result explains the published instruction for this request.

The parser uses the crawler's named group when present, otherwise the wildcard group. It combines repeated groups for the same token, chooses the longest matching rule and prefers Allow when matching rules tie. A wildcard can match a sequence of characters, while a final dollar sign anchors the end. These rules follow the Robots Exclusion Protocol.

Keep crawler purposes separate

The list includes GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, CCBot, Bytespider, Meta-ExternalAgent and Amazonbot. It uses the same reference catalog as the user agent checker.

Some entries collect material for training, some support search and others act on a user's request. Google-Extended and Applebot-Extended are policy controls rather than separate HTTP crawlers. Review the owner's documentation when choosing the activity you want to permit. A training preference and a search indexing preference deserve deliberate, separate decisions.

Review a result before editing policy

Start with the exact hostname and path you intend to make available. Read the matching directive and locate it in the served file. If a host or content delivery service manages robots.txt, use its settings to update the rule. A local file edit will not help if another layer serves a different policy.

After publishing, open robots.txt again and repeat the check for the same path. Keep a copy of the prior and revised policy with the publication date. For actual request outcomes, inspect your trusted server or edge logs. The crawler troubleshooting guide explains how to investigate access failures.

Handle missing files and failed reads

A 404 or 410 response is reported as a missing robots.txt file, with rules marked not specified. A failed request, blocked destination or HTML challenge produces an error instead of a crawler verdict. Correct the address or review the site's availability before trying again.

The fetch is limited to public http and https domains, a ten-second deadline, five redirects and the first 500 KiB of the file. Each redirect destination is checked. Local addresses, private network destinations, custom ports and URLs containing sign-in details are refused. The tool only reads the public file; it does not edit your site or save a monitoring history.

Which hostname should I enter?

Enter the hostname that serves the page you are investigating. The apex domain and a www subdomain can serve different files, as can a documentation subdomain. If you paste a full website URL, the checker uses that origin's robots.txt; set the path field separately for the page you want to evaluate. Follow an intentional redirect to the final policy file and review the checked URL shown with the result.

Use the result in a wider review

This check reads published policy, not actual crawler access or citation eligibility. With that scope clear, use the matching rule as one concrete input to an access investigation. Combine it with the page response and a verified request record when diagnosing a problem.

When access is working, review whether the page explains the customer question accurately. The AI optimization guide connects access, useful content and answer review. SearchSeal's visibility tracker supports recurring answer checks and page fixes if you want to continue from diagnosis into an ongoing improvement workflow.