Live · benchmarked against 10,000 sites

AI Crawlability Checker
26 checks. Fix files out.

ChatGPT, Claude, and Perplexity document separate crawler or retrieval controls, and some sites block one without knowing it. CiteFuel analyzes robots.txt policy for 14 documented crawler/product tokens, compares labeled server responses, validates optional llms.txt file hygiene, applies an internal passage-clarity rubric, and ships reviewable artifacts. External visibility is measured separately.

26 checks 14-token robots matrix 10K-site benchmark ~90s to your report 100% public methodology

The AI crawlability gap — 10,000 sites, June 2026

Google-indexed ≠ AI-citable.

12.5%

Disallow at least one tracked token

In CiteFuel's crawl study of 10,000 Tranco-ranked domains, 12.5% published a robots.txt disallow rule that affected at least one tracked crawler or product-control token. Another 1.9 percentage points came from wildcard rules rather than a named token. This measures published policy, not enforcement.

93.1%

Have no llms.txt at all

Only 6.9% of sampled sites published llms.txt, a voluntary Markdown site-map proposal. This measures adoption only. Google Search ignores llms.txt, and CiteFuel has not established that the file improves rankings or citations on other engines.

2.9%

Published all five measured configuration items

Across five CiteFuel configuration signals — optional llms.txt, Organization schema, crawler policy, canonical tags, HTTPS — only 2.9% of sites hit all five. The median site has 2 of 5 signals addressed. This describes configuration adoption, not ranking, recommendation, or citation.

Source: CiteFuel 10,000-domain crawl study, June 2026 · Methodology

What CiteFuel checks

26 signals. 5 categories.

01 AI Crawler Access

Robots.txt policy analysis across 14 crawler and product tokens: GPTBot, OAI-SearchBot (ChatGPT), ClaudeBot, anthropic-ai (Claude), PerplexityBot, Google-Extended (Gemini), Meta-ExternalAgent, YouBot, Applebot-Extended, and more. Google-Extended and Applebot-Extended are labeled as non-requesting controls. Wildcard fallback analysis catches inherited disallow rules; training, retrieval, user-request, and product-policy intent stay separate.

02 llms.txt Validity

Presence check, 11 syntax/format checks, safe absolute-URL review, and llms-full.txt companion detection. If missing or malformed, CiteFuel can generate an optional proposal-format draft from the audited sitemap. Review every URL before publishing; Google Search ignores llms.txt.

03 Structured-Data Consistency

JSON-LD validation for Organization, WebSite, and Article entities with AI-specific checks: absolute @id URL syntax, applicable types, speakable markup, publisher cross-reference, and breadcrumb structure. External identity URLs are allowed. Valid structured data can help search engines understand visible page content and support eligible rich-result features. It does not establish source credibility or guarantee ranking or citation.

04 Passage-Level Citability

An LLM scores your key passages (0–10) for extractability: are they structured as standalone factual assertions? Do they cite sources, define terms, or answer questions directly? Weak passages are rewritten in the Starter Audit deliverable.

05 Technical Foundation

HTTPS, canonical tags, sitemap validity, response behavior, and PageSpeed metrics (LCP/CLS/INP). Field data is preferred; Lighthouse fallback is labeled as lab evidence. These are technical-quality diagnostics, not proprietary “AI-friendly” ranking thresholds.

06 Live AI Answer Presence

CiteFuel samples live answer engines to measure positive brand-presence readings and, where the provider supplies them, source references. A gap between the internal Readiness score and measured presence is diagnostic evidence, not proof of one specific ranking cause.

Read the full methodology (all 26 checks + weights) →

Why AI crawlability is different from SEO crawlability

Different bots. Different rules.

Traditional SEO crawlability asks one question: can Googlebot reach and index your pages? AI crawlability is a multi-layer problem. At the access layer, ChatGPT and Claude use distinct robots.txt user-agent tokens (GPTBot, ClaudeBot) that are entirely separate from Googlebot — a site can pass Google indexability and fail AI access simultaneously. Many robots.txt files written before 2023 contain wildcard Disallow: / rules that block all unknown crawlers, inadvertently including AI agents added after the rule was written.

Beyond access, CiteFuel separately inventories normal indexability, visible evidence, applicable structured data, content clarity, and independent entity signals. llms.txt is optional and ignored by Google Search; schema does not confer credibility; and a rewritten passage cannot force citation. These checks are workflow diagnostics that should be evaluated alongside actual rankings, citations, reviews, links, and repeated presence measurements.

CiteFuel's AI crawlability checker addresses all three layers: access, structure, and citability. The result is a benchmarked score — calibrated against 10,000 real sites — and a set of fix files you deploy the same day. For a deeper guide to the access layer specifically, see How to configure robots.txt for AI crawlers. For schema markup, see Schema Markup for AI Search.

CiteFuel is also not the only tool in this space. See how we compare to LLMrefs, Rankability, and Sona — honest feature tables, verified June 2026.

Frequently asked questions

What is AI crawlability?

AI crawlability is whether a named crawler or user-triggered fetcher can access and read a website's content. It is distinct from citation: access is necessary for direct retrieval but does not guarantee indexing, ranking, recommendation, or citation. CiteFuel's broader audit separately reports optional-file hygiene, structured data, content heuristics, and measured answer-engine presence.

What does CiteFuel check in its AI crawlability audit?

CiteFuel's 26-check Readiness audit covers five internal categories: crawler and product-token policy; optional llms.txt format and link hygiene; applicable JSON-LD validity; content and entity heuristics; and technical foundations such as HTTPS, canonicals, performance, and freshness. The headline then blends Readiness with a separately sampled positive brand-presence rate. Presence is not uniformly a recommendation or citation. Coverage, unavailable evidence, and confidence are reported separately, and the score is not a ranking guarantee.

Which AI crawlers does the checker test?

CiteFuel evaluates 14 documented crawler and product tokens, including GPTBot and OAI-SearchBot, ClaudeBot and anthropic-ai, PerplexityBot, Google-Extended, Meta-ExternalAgent, YouBot, and Applebot-Extended. Training, retrieval, user-triggered, and control tokens are labeled separately. Google-Extended and Applebot-Extended are non-requesting product-policy tokens; Google-Extended has no effect on Google Search or AI Overviews rankings.

My site is indexed by Google — why would AI crawlability be an issue?

Googlebot and non-Google crawlers use separate robots.txt tokens, so a publisher can choose different access policies. For Google AI Overviews and AI Mode, normal Googlebot indexing and snippet eligibility apply; Google says there are no additional technical requirements, no special AI schema, and llms.txt is ignored. Other engines document their own retrieval controls.

How common are AI crawlability problems?

In CiteFuel's June 2026 crawl frame, 12.5% disallowed at least one tracked crawler token and 1.9% used a wildcard rule that affected tracked agents. The study also measured adoption of optional llms.txt and Organization schema. Those adoption statistics describe technical configuration; they do not prove whether a domain ranks or is cited.

What fix files does CiteFuel generate?

Current paid plans can provide reviewable artifacts such as an optional llms.txt draft built from the audited sitemap, a robots.txt policy suggestion, suggested passage rewrites scored against CiteFuel's editorial rubric, and applicable JSON-LD suggestions. Check current pricing and scope, and verify every URL, fact, policy choice, and live response before deployment.

Check your AI crawlability — free, no card, 90 seconds.

26 published checks. Historical five-signal configuration context kept separate from the current score. Reviewable artifacts generated on the spot.

Check AI crawlability →