AI Crawlability Checker
26 checks. Fix files out.
ChatGPT, Claude, and Perplexity document separate crawler or retrieval controls, and some sites block one without knowing it. CiteFuel analyzes robots.txt policy for 14 documented crawler/product tokens, compares labeled server responses, validates optional llms.txt file hygiene, applies an internal passage-clarity rubric, and ships reviewable artifacts. External visibility is measured separately.
Queued…
The AI crawlability gap — 10,000 sites, June 2026
Google-indexed ≠ AI-citable.
Disallow at least one tracked token
In CiteFuel's crawl study of 10,000 Tranco-ranked domains, 12.5% published a robots.txt disallow rule that affected at least one tracked crawler or product-control token. Another 1.9 percentage points came from wildcard rules rather than a named token. This measures published policy, not enforcement.
Have no llms.txt at all
Only 6.9% of sampled sites published llms.txt, a voluntary Markdown site-map proposal. This measures adoption only. Google Search ignores llms.txt, and CiteFuel has not established that the file improves rankings or citations on other engines.
Published all five measured configuration items
Across five CiteFuel configuration signals — optional llms.txt, Organization schema, crawler policy, canonical tags, HTTPS — only 2.9% of sites hit all five. The median site has 2 of 5 signals addressed. This describes configuration adoption, not ranking, recommendation, or citation.
Source: CiteFuel 10,000-domain crawl study, June 2026 · Methodology
What CiteFuel checks
26 signals. 5 categories.
01 AI Crawler Access
Robots.txt policy analysis across 14 crawler and product tokens: GPTBot, OAI-SearchBot (ChatGPT), ClaudeBot, anthropic-ai (Claude), PerplexityBot, Google-Extended (Gemini), Meta-ExternalAgent, YouBot, Applebot-Extended, and more. Google-Extended and Applebot-Extended are labeled as non-requesting controls. Wildcard fallback analysis catches inherited disallow rules; training, retrieval, user-request, and product-policy intent stay separate.
02 llms.txt Validity
Presence check, 11 syntax/format checks, safe absolute-URL review, and llms-full.txt companion detection. If missing or malformed, CiteFuel can generate an optional proposal-format draft from the audited sitemap. Review every URL before publishing; Google Search ignores llms.txt.
03 Structured-Data Consistency
JSON-LD validation for Organization, WebSite, and Article entities with AI-specific checks: absolute @id URL syntax, applicable types, speakable markup, publisher cross-reference, and breadcrumb structure. External identity URLs are allowed. Valid structured data can help search engines understand visible page content and support eligible rich-result features. It does not establish source credibility or guarantee ranking or citation.
04 Passage-Level Citability
An LLM scores your key passages (0–10) for extractability: are they structured as standalone factual assertions? Do they cite sources, define terms, or answer questions directly? Weak passages are rewritten in the Starter Audit deliverable.
05 Technical Foundation
HTTPS, canonical tags, sitemap validity, response behavior, and PageSpeed metrics (LCP/CLS/INP). Field data is preferred; Lighthouse fallback is labeled as lab evidence. These are technical-quality diagnostics, not proprietary “AI-friendly” ranking thresholds.
06 Live AI Answer Presence
CiteFuel samples live answer engines to measure positive brand-presence readings and, where the provider supplies them, source references. A gap between the internal Readiness score and measured presence is diagnostic evidence, not proof of one specific ranking cause.
Why AI crawlability is different from SEO crawlability
Different bots. Different rules.
Traditional SEO crawlability asks one question: can Googlebot reach and index your pages? AI
crawlability is a multi-layer problem. At the access layer, ChatGPT and Claude use distinct
robots.txt user-agent tokens (GPTBot, ClaudeBot) that are entirely separate from Googlebot —
a site can pass Google indexability and fail AI access simultaneously. Many robots.txt files
written before 2023 contain wildcard Disallow: / rules that block all unknown
crawlers, inadvertently including AI agents added after the rule was written.
Beyond access, CiteFuel separately inventories normal indexability, visible evidence, applicable structured data, content clarity, and independent entity signals. llms.txt is optional and ignored by Google Search; schema does not confer credibility; and a rewritten passage cannot force citation. These checks are workflow diagnostics that should be evaluated alongside actual rankings, citations, reviews, links, and repeated presence measurements.
CiteFuel's AI crawlability checker addresses all three layers: access, structure, and citability. The result is a benchmarked score — calibrated against 10,000 real sites — and a set of fix files you deploy the same day. For a deeper guide to the access layer specifically, see How to configure robots.txt for AI crawlers. For schema markup, see Schema Markup for AI Search.
CiteFuel is also not the only tool in this space. See how we compare to LLMrefs, Rankability, and Sona — honest feature tables, verified June 2026.
Frequently asked questions
What is AI crawlability?
AI crawlability is whether a named crawler or user-triggered fetcher can access and read a website's content. It is distinct from citation: access is necessary for direct retrieval but does not guarantee indexing, ranking, recommendation, or citation. CiteFuel's broader audit separately reports optional-file hygiene, structured data, content heuristics, and measured answer-engine presence.
What does CiteFuel check in its AI crawlability audit?
CiteFuel's 26-check Readiness audit covers five internal categories: crawler and product-token policy; optional llms.txt format and link hygiene; applicable JSON-LD validity; content and entity heuristics; and technical foundations such as HTTPS, canonicals, performance, and freshness. The headline then blends Readiness with a separately sampled positive brand-presence rate. Presence is not uniformly a recommendation or citation. Coverage, unavailable evidence, and confidence are reported separately, and the score is not a ranking guarantee.
Which AI crawlers does the checker test?
CiteFuel evaluates 14 documented crawler and product tokens, including GPTBot and OAI-SearchBot, ClaudeBot and anthropic-ai, PerplexityBot, Google-Extended, Meta-ExternalAgent, YouBot, and Applebot-Extended. Training, retrieval, user-triggered, and control tokens are labeled separately. Google-Extended and Applebot-Extended are non-requesting product-policy tokens; Google-Extended has no effect on Google Search or AI Overviews rankings.
My site is indexed by Google — why would AI crawlability be an issue?
Googlebot and non-Google crawlers use separate robots.txt tokens, so a publisher can choose different access policies. For Google AI Overviews and AI Mode, normal Googlebot indexing and snippet eligibility apply; Google says there are no additional technical requirements, no special AI schema, and llms.txt is ignored. Other engines document their own retrieval controls.
How common are AI crawlability problems?
In CiteFuel's June 2026 crawl frame, 12.5% disallowed at least one tracked crawler token and 1.9% used a wildcard rule that affected tracked agents. The study also measured adoption of optional llms.txt and Organization schema. Those adoption statistics describe technical configuration; they do not prove whether a domain ranks or is cited.
What fix files does CiteFuel generate?
Current paid plans can provide reviewable artifacts such as an optional llms.txt draft built from the audited sitemap, a robots.txt policy suggestion, suggested passage rewrites scored against CiteFuel's editorial rubric, and applicable JSON-LD suggestions. Check current pricing and scope, and verify every URL, fact, policy choice, and live response before deployment.
Related tools & reading
llms.txt Checker
Validate or generate your llms.txt against 11 spec checks. Only 6.9% of sites have one.
AI Crawler Access Checker
robots.txt matrix for all 14 AI user-agents — with the fix block to paste in.
Does Your Site Appear in ChatGPT?
How to test AI citation presence and what the gaps mean for your traffic.
Check your AI crawlability — free, no card, 90 seconds.
26 published checks. Historical five-signal configuration context kept separate from the current score. Reviewable artifacts generated on the spot.
Check AI crawlability →