Free robots.txt checker for search and AI crawlers

Your robots.txt decides which crawlers may read your site, including the ones behind ChatGPT, Claude, Perplexity and Gemini. See what yours allows, what it blocks, and what to change.

Free. No signup, no card, no email.

What is the Sophyx Robots.txt Checker?

A free tool that reads your robots.txt, the small text file that tells crawlers where they may go. It shows which search and AI crawlers you allow or block, whether your sitemap is listed, and which rules might hide pages you want found.

If a crawler can't read a page, that page can't help you show up in search or in AI answers.

Example robots.txt check

Here is a check of a made-up bike shop's robots.txt. Two AI crawlers are blocked, probably by accident.

northshorebikes.ca/robots.txt

Example

2 AI crawlers blocked, sitemap missing

  • GPTBot

    OpenAI's crawler that gathers pages for training its models.

    Allowed
  • OAI-SearchBot

    Lets ChatGPT search find and link your pages.

    Allowed
  • ClaudeBot

    Disallow: / keeps Anthropic's crawler out of the whole site.

    Blocked
  • PerplexityBot

    Lets Perplexity find and cite your pages.

    Allowed
  • Google-Extended

    Stops Gemini apps from using your pages. Google Search and AI Overviews are not affected.

    Blocked
  • Sitemap

    No Sitemap line, so crawlers have to find your pages on their own.

    Missing
  • Googlebot and Bingbot

    Search crawling works as normal.

    Allowed

The rules behind this result

User-agent: *
Disallow: /cart/
Disallow: /checkout/

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

What to fix

  1. 1.Decide whether you want Claude and Gemini to use your pages. If you do, delete the two Disallow: / blocks.
  2. 2.Add Sitemap: https://northshorebikes.ca/sitemap.xml at the end of the file.
  3. 3.Keep /cart/ and /checkout/ blocked. They don't belong in search or AI answers.
Example only. North Shore Bike Co. (1525 Lonsdale Ave, North Vancouver) is a made-up shop.

Blocking a crawler can be the right choice. The point is to choose it on purpose, knowing what each one does.

How it works

  1. Step 1

    Enter your website

    We look for robots.txt at the root of your domain.

  2. Step 2

    We read the rules

    User-agent, Allow, Disallow and Sitemap lines, for search and AI crawlers.

  3. Step 3

    Get a plain report

    What each crawler can reach, what looks risky, and what to change.

What you get

Is the file there?

Whether robots.txt exists at the root of your domain and loads.

Who gets in

Which search and AI crawlers are allowed or blocked: GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Googlebot and Bingbot.

Is your sitemap listed?

Whether crawlers are pointed to the list of your pages.

What to change

Rules that block too much, old staging rules, and the areas worth keeping blocked.

Who it's for

  • Business owners who want their site found by search and AI.
  • Marketing teams checking that key pages can be crawled.
  • SEO and GEO consultants doing technical audits.
  • Developers checking rules after a launch, move or redesign.

What it often finds

  • AI crawlers blocked without anyone deciding to.
  • A Disallow: / left over from a staging site.
  • Blog or service pages blocked by a broad rule.
  • No sitemap listed.

robots.txt, explained

What is a robots.txt file?

A plain text file at yourdomain.com/robots.txt with rules for crawlers. User-agent says which crawler a rule is for. Disallow says where it may not go. Sitemap points to the list of your pages.

It is simple, but one wrong line can hide your whole site.

Which AI crawlers should you know?

  • GPTBot: OpenAI's crawler for training its models
  • OAI-SearchBot: lets ChatGPT search show and link your pages
  • ClaudeBot: Anthropic's crawler for Claude
  • PerplexityBot: lets Perplexity find and cite your pages
  • Google-Extended: controls whether Gemini apps may use your pages; Google Search is not affected

You can allow some and block others. Decide on purpose.

What should robots.txt include?

  • Blocks for admin, cart, checkout and internal search pages
  • Blocks for staging or test areas
  • A Sitemap line
  • Rules for specific crawlers, only where you mean them

Block what shouldn't be found and leave the rest open.

What should you avoid?

  • Disallow: / for all crawlers
  • Blocking your blog, services or products
  • Blocking the CSS or JavaScript your pages need to display
  • Old rules copied from a previous site
  • Using robots.txt to hide private information

Does robots.txt control indexing?

Not exactly. It controls crawling. A blocked page can still show up in search if other sites link to it, just without its content.

To keep something private, put it behind a login or use a noindex tag. robots.txt isn't a lock.

After you get your robots.txt report

  1. 1.Decide which AI crawlers you want to allow, and make the rules match.
  2. 2.Remove any broad block that hides pages you want found.
  3. 3.Add a Sitemap line if it's missing.
  4. 4.Keep admin, cart, checkout and internal search pages blocked.
  5. 5.Don't rely on robots.txt to protect private information.
  6. 6.Run the check again after you change the file.

Next: see what AI says about you

Once crawlers can reach your pages, the next question is what AI does with them. Run the free check to see whether ChatGPT and Gemini name you, and what they get wrong.

More free tools

Check your visibility

Fix your site

Explore Sophyx

See all free tools

Questions people ask

Is the Sophyx Robots.txt Checker free?
Yes. Check your robots.txt free at app.sophyx.io/robot-txt-checker. No signup, no card, no email.
Do I need to create an account?
No. It works without a Sophyx account.
What is a robots.txt file?
A plain text file that gives crawl rules to search engines and other well-behaved bots. It lives at https://yourdomain.com/robots.txt.
What does the tool check?
Whether the file exists and loads, which paths are allowed or blocked, which search and AI crawlers can get in, whether a sitemap is listed, and which rules look risky.
Can robots.txt improve SEO?
It helps crawlers skip low-value areas and reach the pages that matter. It doesn't guarantee rankings, indexing or traffic.
Can robots.txt block my website from Google?
Yes, a broad rule can. Disallow: / under User-agent: * tells most crawlers to stay out of the whole site.
Does robots.txt remove pages from search results?
Not reliably. It controls crawling, not removal. To remove or hide a page, use a noindex tag or put it behind a login.
Should I include my sitemap in robots.txt?
Usually, yes. A Sitemap line helps crawlers find your important pages.
Can I use this for AI crawlers?
Yes. It shows whether GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended are allowed. It can't promise how any AI system will use your pages once it can read them.
How often should I check robots.txt?
After launches, redesigns, site moves and CMS changes, and any time you edit the file.
How does robots.txt relate to AI SEO versus traditional SEO?
Both start with crawlers reaching your pages. Google and Bing need it to rank you, and AI tools need it to read you. One blocked folder can hurt both.
Does robots.txt protect data privacy and sensitive pages?
No. It is a request that well-behaved crawlers follow, not a lock, and anyone can still open a blocked page. Put anything private behind a login.
What is Answer Engine Optimization (AEO) and how does this tool help?
AEO means making your business easy for AI assistants such as ChatGPT, Gemini, Claude and Perplexity to find, understand and name in their answers. This tool checks whether search and AI crawlers can reach the pages you want found. It is one fix you can make today. To see what AI actually says about you, run the free check at app.sophyx.io/brandanalysis.
How is AI SEO different from traditional SEO?
SEO is about ranking in a list of links on Google and Bing. AI SEO (also called AEO or GEO) is about what the AI says when it answers: whether it names you, and whether it gets your facts right. AI builds that answer from everything the web says about you, so clear pages, labelled facts and mentions on other sites all count. The same work helps both.
Which AI engines and platforms does Sophyx optimize for?
The fixes help on every assistant, because ChatGPT, Gemini, Claude, Perplexity, Copilot and Google AI Overviews all read the same public web. What we track is narrower: ChatGPT and Gemini, including Google AI Overviews. Your own website matters most, and sites like LinkedIn, Reddit and review sites also shape what AI says about you.
How does this free tool connect to Sophyx's AI visibility platform?
This tool checks or fixes one thing. Sophyx does the rest: it asks ChatGPT and Gemini the questions your customers ask, shows you where you are missing or described wrong, writes the fixes, and publishes them on your own website once you approve. Start with the free check at app.sophyx.io/brandanalysis.
Does Sophyx offer help for marketing agencies?
Yes. Through the Agency Partner Program, agencies run Sophyx for every client under their own brand, with white-label, client-ready reports and priority support. Many agencies use these free tools for a quick first look at a client's site.

Check your robots.txt for free

See which search and AI crawlers can read your site, and what is blocked by mistake. Free, no signup.

Check my robots.txt
Book a Demo