Skip to content

Free GEO audit: how search engines and AI see your page

One URL, one report: can Google and AI search reach, index and understand the page, what to fix first, and exactly how, with the official source behind every check.

URL mode fetches one public page plus its robots.txt, llms.txt and sitemap through our Worker; the analysis runs in your browser. Pasted HTML and robots.txt drafts never leave your browser.Bookmarklet for pages behind a login.

How to use the audit

  1. Enter a URL (any page on your site), or paste the page HTML or a robots.txt draft. URL mode fetches the page once, with its redirects, headers, robots.txt, llms.txt and sitemap; the analysis itself runs in your browser.
  2. Start with "Fix these first". Fixes cost indexing or visibility; improvements are documented best practice. Open any item for why it matters, how to fix it, the evidence from your page, a copy-paste snippet where one applies, and the official source.
  3. Check the previews. See how the title and description can render in results on desktop and phone, and how the link looks when shared in chat and social apps.
  4. Work through the areas: indexability, search appearance, content and HTML (with the page’s heading outline), links (press "Check link status" to test up to 15 of them, the canonical target first), structured data (every JSON-LD item as a readable tree, checked against Google’s rich result requirements) and AI & GEO. The bar under the header jumps between them.
  5. Scan the whole site. Use "Scan a site": drag the button to your bookmarks bar, open your site and click it. Up to 50 pages from your sitemap are checked in your browser for duplicate titles and descriptions, missing headings, noindex, error or redirecting URLs in the sitemap, canonicals that point elsewhere, and broken internal links. Export the page table as CSV.
  6. Measure speed where the report offers it: real-user Core Web Vitals from the Chrome UX Report, rated with Google’s published thresholds, and a Lighthouse lab test through PageSpeed Insights.
  7. Crawler by crawler. Crawlers are grouped by what they do: classic search, AI search, user-triggered fetches, training and datasets. A training block you chose is shown as policy, not as a problem.
  8. Pick your policy. The default, "Keep my current rules", reports what is true today. Choose another preset to see what would change; the table, the fixes and the proposed file recompute instantly, with no refetch. Use "Custom" to set crawlers one by one.
  9. Test paths. Add the URLs that matter (a product page, a blog folder) to see which crawlers can reach each one and which line decides it.
  10. Ship the file. Copy or download the proposed robots.txt. Check the diff first: only the listed crawlers change.
  11. Optional: run the live probe to catch firewall or CDN rules that refuse a crawler robots.txt allows, and check your llms.txt links.

Worked examples

1. A deliberate "allow the citers, block the costly crawlers" policy

PeptideClock’s robots.txt (read 2026-09-23) allows everyone under User-agent: *, names 18 AI crawlers in an explicit Allow: / group, and disallows three: Bytespider, meta-externalagent and Amazonbot, with a comment explaining they have "high crawl cost, no referral and no citation path".

The checker reports no fixes for search and AI search: Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot and PerplexityBot can all read the calculator page, and the three blocks appear as policy. The default "Keep my current rules" judges nothing else. Switch to "AI search yes, AI training no" and the tool lists GPTBot, ClaudeBot and CCBot as not matching, because this site allows training crawlers on purpose, and proposes the file that would block them.

2. Cloudflare’s managed robots.txt contradicts your own file

When the Cloudflare setting "Set your preference to block training in robots.txt" is on, Cloudflare "will prepend our managed robots.txt before your existing robots.txt", between # BEGIN Cloudflare Managed content and # END Cloudflare Managed Content, with Disallow: / groups for crawlers such as GPTBot, ClaudeBot, Google-Extended and CCBot, and a Content-signal: search=yes, ai-train=no line (Cloudflare docs, read 2026-09-23).

If your own file also has User-agent: GPTBot / Allow: /, GPTBot is now named in two groups. Under RFC 9309 the groups merge and the tie between Allow: / and Disallow: / goes to Allow; a crawler whose parser reads only the first group sees the managed Disallow. The checker shows both outcomes, the line numbers, and leaves the managed block out of the proposed file because Cloudflare adds it at the edge.

3. robots.txt allows a crawler, the firewall does not

Illustrative case: robots.txt allows GPTBot, but a WAF rule matching its user-agent returns 403 while browsers get 200. robots.txt checks cannot see this. The live probe requests the page once as a browser and once per crawler user-agent and reports the difference, labelled with its limit: it comes from our Worker, not from the vendor’s published IP ranges, so rules that verify crawler IPs may treat the real crawler differently.

Crawler reference

Every row comes from the vendor’s own crawler page, with the date we read it. When a vendor documents nothing, the row says so.

AI and search crawler reference
User-agent tokenPurposeWhat it feedsrobots.txtVendor source
GooglebotGoogle
Classic search
Google Search, including Discover and all Google Search features
Follows robots.txt
developers.google.comretrieved 2026-09-23
BingbotMicrosoft
Classic search
Bing search index
Follows robots.txt
www.bing.comretrieved 2026-09-23
ApplebotApple
Classic search
Spotlight, Siri and Safari search results
Follows robots.txt
support.apple.comretrieved 2026-09-23
OAI-SearchBotOpenAI
AI search
ChatGPT search results
Follows robots.txt
developers.openai.comretrieved 2026-09-23
Claude-SearchBotAnthropic
AI search
Claude search result quality
Follows robots.txt
support.claude.comretrieved 2026-09-23
PerplexityBotPerplexity
AI search
Perplexity search results
Follows robots.txt
docs.perplexity.airetrieved 2026-09-23
Amzn-SearchBotAmazon
AI search
Search experiences in Amazon products such as Alexa
Follows robots.txt
developer.amazon.comretrieved 2026-09-23
Meta-WebIndexerMeta
AI search
Meta AI search results and citations
Follows robots.txt
developers.facebook.comretrieved 2026-09-23
ChatGPT-UserOpenAI
User-triggered fetch
Pages ChatGPT opens when a user asks
May ignore robots.txt (user-triggered)
developers.openai.comretrieved 2026-09-23
Claude-UserAnthropic
User-triggered fetch
Pages Claude opens when a user asks
Follows robots.txt
support.claude.comretrieved 2026-09-23
Perplexity-UserPerplexity
User-triggered fetch
Pages Perplexity opens when a user asks
May ignore robots.txt (user-triggered)
docs.perplexity.airetrieved 2026-09-23
Meta-ExternalFetcherMeta
User-triggered fetch
Links fetched at a user request, agentic AI features
May ignore robots.txt (user-triggered)
developers.facebook.comretrieved 2026-09-23
Amzn-UserAmazon
User-triggered fetch
Live answers for Alexa queries
May ignore robots.txt (user-triggered)
developer.amazon.comretrieved 2026-09-23
GPTBotOpenAI
AI training
Training of OpenAI generative AI foundation models
Follows robots.txt
developers.openai.comretrieved 2026-09-23
ClaudeBotalso anthropic-aiAnthropic
AI training
Anthropic model training datasets
Follows robots.txt
support.claude.comretrieved 2026-09-23
Google-ExtendedGoogle
AI training
Gemini model training and grounding in Gemini Apps / Vertex AI
Control token only (no separate crawler)
developers.google.comretrieved 2026-09-23
Applebot-ExtendedApple
AI training
Training of Apple's foundation models (Apple Intelligence)
Control token only (no separate crawler)
support.apple.comretrieved 2026-09-23
Meta-ExternalAgentMeta
AI training
Training foundation AI models and product indexing
Follows robots.txt
developers.facebook.comretrieved 2026-09-23
AmazonbotAmazon
AI training
Amazon products and services; may be used to train Amazon AI models
Follows robots.txt
developer.amazon.comretrieved 2026-09-23
CCBotCommon Crawl
Open crawl dataset
Common Crawl open web archive, reused by third parties
Follows robots.txt
commoncrawl.orgretrieved 2026-09-23
BytespiderByteDance
Undocumented
Unknown — no vendor documentation found
Unknown
None found

Methodology and limits

  • robots.txt semantics follow RFC 9309 as Google implements it: groups naming the same crawler are merged; the * group applies only when no group names the crawler; the longest matching path wins and Allow wins a tie; * and $ wildcards; empty rules are ignored (RFC 9309, Google’s interpretation, read 2026-09-23). Our test suite includes Google’s published matching and precedence examples.
  • Vendor-specific fallbacks are applied only where the vendor documents one — for example, Apple says Applebot follows Googlebot rules when it is not named (Apple).
  • Severity is purpose-aware. "Fix" means a crawler that feeds classic search or AI search cannot read the page, or a header removes it from indexes. Blocks of training, dataset and undocumented crawlers are "Policy". User-triggered fetchers are "Check".
  • Headers count. An X-Robots-Tag: noindex header or meta robots noindex overrides a permissive robots.txt for indexing crawlers; both are read.
  • What robots.txt cannot tell you: whether a crawler actually obeys it, whether a firewall blocks it, or whether an engine will cite you. The probe covers the second only partly (see its note). Nothing here measures citations.
  • Every SEO check cites its source: Google Search Central, a web standard, or accessibility guidance. Checks that are sensible but not documented as ranking factors are labelled heuristics. There is no score, and no character limits Google does not publish. FAQ schema is not recommended: Google removed the FAQ rich result documentation on 15 June 2026 because the feature is no longer shown (Search Central changelog).
  • Pages that do not load (404, 5xx, a login, a redirect loop) still get a report of what we saw: the status, every redirect hop and the response headers. Nothing is said about content we never received. A 403 or 429 to our checker is not treated as proof that Google is blocked, because bot protection often treats unknown fetchers differently.
  • The site scan reads pages the way a browser does, from your site’s own origin with cookies left out. A browser can check the status of links on your own site, but for other sites it only sees whether a connection was made, so external 404s are not reported, and on an https site it may not request http:// links at all. A 401, 403 or 429 is listed for you to check by hand, not called broken: logins, bot protection and rate limits often answer a quick series of requests differently. URLs on another host (for example www versus no www) are counted but not read; run the scan from that host to include them. If your site sends Cross-Origin-Opener-Policy: same-origin, the scan still runs, and you move the results over with Download results and Import.
  • Limits: one page per URL check, robots.txt and llms.txt up to 100 KB, a 25-URL sitemap sample, 20 URL checks per hour, 5 probes per hour, 10 link checks and 10 speed checks per hour per IP. The site scan reads up to 50 pages, checks up to 200 internal and 30 external link targets, and has no hourly limit because it runs in your browser.

Troubleshooting

The live robots.txt is not the file in my repository
Look for a Cloudflare managed block (lines starting # BEGIN Cloudflare Managed content). The setting lives in the Cloudflare dashboard under Security → Settings, filtered by Bot traffic: "Set your preference to block training in robots.txt". A CDN can also cache an old file — purge /robots.txt after changing it.
I changed robots.txt but a crawler still behaves the old way
Crawlers cache the file. OpenAI says it can take about 24 hours for its systems to adjust; Meta says to allow up to 24 hours; Google generally caches robots.txt for up to 24 hours (OpenAI, Meta, Google).
robots.txt allows a crawler but my logs show 403s for it
A firewall, bot-management or "block AI bots" setting is answering its user-agent. Run the live probe, then check user-agent rules in your WAF. Perplexity documents allow-listing its crawlers by user-agent plus its published IP ranges (Perplexity).
The checker says "Blocked" for a path I thought was open
Read the deciding line in the table. Common causes: a longer Disallow (/blog/ beats /), a wildcard (/*? blocks every URL with a query string), or a named group that replaces the * group entirely — a crawler with its own group ignores the * rules.

Questions

Why is there no SEO score?

Because no public formula would be honest. Google does not publish one, and a single number hides what matters: a noindex tag or a blocked Googlebot outweighs every other check combined. The audit instead counts, per area, what to fix and what to improve, ranks the fixes, and links the official source for each rule.

Why doesn’t it warn that my title or meta description is too long?

Google publishes no character limit for either. Its documentation says there is no limit on length and that results are truncated "as needed, typically to fit the device width" (title link and snippet docs). The audit therefore previews how your title renders at desktop and phone width and tells you where it would be cut, which is the part you can act on.

What does the AI & GEO section actually check?

What is documented. Google says a page must be indexed and eligible for a snippet to appear in its generative AI features, so noindex, nosnippet and a blocked Googlebot are reported as fixes. AI search products document their own search crawlers, so the audit checks those against your robots.txt. Signals that are sensible but not documented as ranking factors, such as a visible author or date, are labelled heuristics.

If I block GPTBot, will my site disappear from ChatGPT?

No. OpenAI documents two separate controls: OAI-SearchBot decides whether you can appear in ChatGPT search answers, and GPTBot covers model training. "Each setting is independent of the others", so blocking GPTBot while allowing OAI-SearchBot keeps you eligible for ChatGPT search (OpenAI crawler docs, read 2026-09-23).

Does blocking Google-Extended remove me from Google Search or AI Overviews?

Google says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal". It controls use in Gemini training and grounding. Google Search features, including AI Overviews, run on the Googlebot crawl, so blocking Googlebot is what removes a page (Google common crawlers doc, read 2026-09-23).

Why does the tool call my Bytespider block "policy" instead of an error?

We found no first-party documentation for Bytespider, so its purpose is unknown and blocking it costs no visibility this tool can measure. Blocks of training, dataset or undocumented crawlers are treated as your policy. Only blocks that cost search or AI-search visibility you asked for are marked "Fix".

Do AI assistants obey robots.txt when a user asks them to open my page?

It depends on the vendor. OpenAI says robots.txt rules "may not apply" to ChatGPT-User; Perplexity says Perplexity-User "generally ignores robots.txt rules"; Meta says Meta-ExternalFetcher "may bypass robots.txt rules". Anthropic says its bots respect robots.txt and lists Claude-User among them. The reference table below carries each vendor’s wording and source.

Does the site scan send my pages to your server?

No. The scanner runs in your browser, on your own site: it reads your robots.txt, sitemap and pages from your own connection, without cookies, and sends only the findings to the pyralislabs.io tab you have open. Our Worker is not involved and makes no requests to your site. You can also download the results as a file and import them later.

Why does the site scan use a bookmarklet?

Browsers do not let one website read another (the same-origin policy), so a page on pyralislabs.io cannot fetch your pages. Running the scanner on your own site makes every request same-origin, the same as you opening those pages. If your site’s Content-Security-Policy blocks the scanner script, the report tab gives you the same code to paste into the browser console instead.

Do I need an llms.txt file?

It is optional. llms.txt is a proposed convention (llmstxt.org) for giving agents a curated, markdown map of a site. Google states the file is not needed for Google Search and does not affect rankings (Search Central changelog, 15 June 2026). If you publish one, this tool checks its structure and links; it never overwrites a file you already have.

Will the proposed robots.txt delete my existing rules?

No. The builder rewrites only the AI crawlers whose current verdict differs from the policy you picked. Every other line — comments, sitemaps, search-engine groups, private-path rules — is kept, and the diff view shows exactly which lines change. When it opens a crawler, it carries over your private-path Disallow rules.

Can I test a robots.txt before I deploy it?

Yes. Use "Paste robots.txt". The draft is parsed in your browser only; nothing is sent to us. Add a page path to see whether a specific URL is reachable for each crawler.