# robots.txt Tester for AI Crawlers

> Free robots.txt tester built for AI crawlers. See which of GPTBot, ClaudeBot, PerplexityBot, Googlebot and Bingbot you allow, plus a CDN firewall test.

This is the Markdown copy of https://usesuperflow.ai/tools/robots-txt-ai-checker, published for AI agents and scripts.

## What it does

Test your robots.txt against GPTBot, ClaudeBot, PerplexityBot, Googlebot, and every other crawler that decides whether AI can cite you.

## How to use it

1. **Paste your URL.** We fetch the robots.txt at your domain root and parse it exactly the way a crawler does.
2. **We test every AI crawler.** Eleven answer engines and six training crawlers, each evaluated against the path you gave us, with the deciding rule shown.
3. **We test your firewall too.** We request your page as GPTBot and compare it to a browser request, because a CDN can block crawlers your robots.txt welcomes.

## Call it directly

**HTTP:** `POST https://usesuperflow.ai/api/tools/robots-txt-ai-checker`. No API key, no account.

```bash
curl -sS https://usesuperflow.ai/api/tools/robots-txt-ai-checker \
  -H 'Content-Type: application/json' \
  -d '{"url":"example.com"}'
```

Returns `{ ok, report: { accessScore, crawlers[] with the rule that decided each verdict, firewall, findings[] }, cached, ageSeconds }`.

**MCP:** this tool is `check_robots_txt_for_ai` on Superflow's MCP server at `https://usesuperflow.ai/api/mcp` (Streamable HTTP, no authentication). Setup for every client: https://usesuperflow.ai/tools/mcp

Limits: 10 runs per hour per IP. Allow up to 75 seconds for a response. Results are cached for 24 hours per URL; pass `"refresh": true` to run again.

## Facts

| | |
| --- | --- |
| Cost | Free. No login, no email, no ads. |
| Where it runs | On our server. robots.txt is fetched and parsed, then the page itself is requested with each crawler's real user agent, because a CDN-level block stops a crawler before robots.txt is ever read. |
| API | POST /api/tools/robots-txt-ai-checker with a JSON body of {"url": "example.com"}. Returns the same report the page shows. |
| Rate limit | 10 runs per hour per IP. |
| How long a run takes | Up to 75 seconds. A slow run answers with { status: "pending", runId } instead of a result: call again with just that runId to collect it. Collecting costs no rate-limit slot. |
| Stored data | The report, cached for 24 hours keyed on the URL. Nothing beyond that cache. |
| What it checks | Every AI and search crawler that matters (GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, CCBot, Applebot-Extended, Googlebot, Bingbot and the rest) against the site's robots.txt, reporting which are allowed, which are blocked, and the exact rule that decided each verdict. Plus a live firewall test per crawler. |
| Why the firewall test matters | robots.txt is a request, not a gate. A CDN rule can return 403 to an AI crawler before robots.txt is read, so a site whose robots.txt says yes can still be invisible. This is the check almost nothing else runs. |
| Related tool | This is the access-scoped view. Use the AI Visibility Checker for the whole-page verdict, including readability and structure. |

## Questions

### How do I test if my robots.txt blocks GPTBot?

Paste your URL above. We fetch your robots.txt, parse it the way a real crawler does, and evaluate every AI user agent against the exact path you gave us. The results table shows each crawler, whether it is allowed or blocked, and which rule decided it.

### Why does my robots.txt look fine but AI still cannot read my site?

Almost always a firewall. Cloudflare and other CDNs ship one-click toggles that block AI crawlers at the edge, before the request ever reaches your server or your robots.txt. We test for this directly by requesting your page twice, once as a browser and once as GPTBot, and comparing the responses. Nothing in your CMS will show you this.

### What does Disallow: / actually block?

Everything on the site, for whichever user agent group it appears under. The subtlety is that a crawler follows exactly one group, the one whose User-agent token is the longest match for its name. So a Disallow: / under User-agent: * does not apply to GPTBot if there is also a User-agent: GPTBot group anywhere in the file, even an empty one.

### Does Allow beat Disallow?

Only when it is at least as specific. The longest matching path pattern wins regardless of the order the rules appear in, and when an Allow and a Disallow match with equal length, Allow wins. This is why adding Allow: / at the bottom of a file that starts with Disallow: / does unblock the site, and it is the rule most robots.txt checkers get backwards.

### Should I block AI crawlers in robots.txt?

It depends which ones. Blocking CCBot, Google-Extended, or Applebot-Extended keeps your content out of model training and costs you nothing in AI answers. Blocking OAI-SearchBot, ChatGPT-User, PerplexityBot, or Claude-SearchBot removes you from the answers themselves. Our table splits the two so you can make that call deliberately.

### Where does robots.txt have to live?

At the root of the domain, at /robots.txt exactly. Crawlers do not look anywhere else, and a robots.txt in a subdirectory does nothing. Each subdomain needs its own, so blog.example.com is not covered by the file at example.com.

## Related tools

- [AI Visibility Checker](https://usesuperflow.ai/tools/ai-visibility-checker) — See whether ChatGPT, Claude, and Perplexity can actually read your site. Markdown copy: https://usesuperflow.ai/tools/ai-visibility-checker.md
- [llms.txt Generator](https://usesuperflow.ai/tools/llms-txt-generator) — Generate a spec-correct llms.txt and llms-full.txt for any site. Markdown copy: https://usesuperflow.ai/tools/llms-txt-generator.md
- [Time Zone Converter & Meeting Planner](https://usesuperflow.ai/tools/meeting-planner) — Convert local times and find overlapping working hours by city or country. Markdown copy: https://usesuperflow.ai/tools/meeting-planner.md
- [Markdown for Agents](https://usesuperflow.ai/tools/markdown-for-agents) — Turn your pages into clean Markdown you can host for AI agents to read. Markdown copy: https://usesuperflow.ai/tools/markdown-for-agents.md

## About

robots.txt Tester for AI Crawlers is one of a set of free tools published by Superflow at https://usesuperflow.ai/tools. None of them require a login, an email address, or payment.

Category: ai-visibility.

Superflow is a website and creative-asset review tool. Its agents watch every page of a site and report what changed. See https://usesuperflow.ai.
