The State of AI Crawlability 2026:
We Crawled 10,000 Popularity-Ranked Domains
Only 6.9% of 10,000 popularity-ranked domains (Tranco ranks 1,001–11,000) publish an llms.txt file. 12.5% disallow at least one tracked crawler or product-control token — and 1.9% do so through a wildcard rather than a named rule. This is a point-in-time configuration study; it did not measure rankings, citations, recommendations, or traffic.
- 1 Only 6.9% of 10,000 popularity-ranked domains at Tranco ranks 1,001–11,000 publish an llms.txt file, according to CiteFuel's June 2026 crawl.
- 2 One in eight reachable websites (12.5%) disallowed at least one tracked crawler or product token in robots.txt. This is a policy measurement, not proof of product-level invisibility.
- 3 1.9% of reachable sites used a wildcard
Disallow: /rule that affected tracked agents without naming each token explicitly. - 4 Only 2.9% of reachable sites score 5 out of 5 on CiteFuel's configuration index. The most common score is 2/5, and 47 sites score 0/5; the index does not measure external visibility.
- 5 Just 18.9% of reachable sites publish Organization schema. Its absence is a configuration observation, not a ranking or citation failure.
- 6 llms.txt adoption is not highest among the biggest websites: it peaks at ranks 2,001–3,000 (9.1%) and is flat or declining among the very top 1,000 sites. Fast-moving mid-tier sites are the early adopters.
- 7 Sites with an llms.txt file had a 20% faster median TTFB in this sample (203ms vs. 255ms). This is a correlational association that may reflect confounding factors; it does not show that llms.txt caused faster responses.
- 8 GPTBot had the highest disallow rate among the tracked tokens (11.3%), followed by ClaudeBot (10.3%), Google-Extended (9.5%), anthropic-ai (6.4%), and PerplexityBot (6.2%). Google-Extended is a non-requesting product-control token, not a crawler user agent.
The five-signal configuration distribution is wide
In June 2026, CiteFuel crawled 10,000 domains from Tranco ranks 1,001–11,000 — a large popularity-ranked cohort, not a statistically representative sample of the open web — to inventory five crawl and page-configuration signals. Of the 10,000 domains queried, 7,543 returned parseable homepages. The rest timed out, returned non-HTML roots, or disallowed our crawler (217 domains honored; see Methodology).
693 of ~10,000 domains have any AI-guidance file
crawler and product-control tokens combined
216 sites publish all five measured items; mode is 2/5
4 in 5 sites lack a structured entity anchor
The measured conclusion is narrower: most sites did not publish every item in CiteFuel's five-signal configuration checklist. This crawl did not measure rankings, recommendations, citations, or traffic, so a low checklist score is not proof that a site is invisible to AI.
This historical dataset provides a reproducible configuration benchmark. The current free CiteFuel audit uses a broader, separately documented methodology and should not be treated as this five-item index.
Check Your Site Free →Chart 1: Five-signal configuration index (0–5)
Each reachable site was scored on five binary signals: llms.txt present,
no disallow rule for the tracked crawler/product tokens, Organization schema present, canonical tag present, and HTTPS redirect active.
Scores range from 0 (none) to 5 (all five).
"Only 2.9% of reachable sites in CiteFuel's June 2026 sample published all five items in its configuration index. The most common score was 2 out of 5. The index did not measure external search or answer-engine visibility."
¶ cite this findingChart 2: Per-token robots.txt disallow rates
We checked every domain's robots.txt against five crawler or product-control tokens using effective-rule
semantics — including wildcard fallback. A token is counted as disallowed when the effective rule is
Disallow: / for that user-agent, whether the token is named explicitly or inherited
from a wildcard group.
Google-Extended is different from the request-capable crawlers in this chart: Google
documents it as a standalone product token that does not make HTTP requests. Its setting does not affect
inclusion or ranking in Google Search, including AI Overviews and AI Mode.
The silent blocking problem
Across the sample, 1.9 percentage points had a wildcard rule that disallowed tracked
tokens without naming those tokens explicitly. For example, GPTBot can inherit
User-agent: * and Disallow: / even when the string "GPTBot" does not appear
in the file. This is an effective-rule observation, not evidence about the administrator's intent.
"In CiteFuel's June 2026 sample, 1.9% of domains had a wildcard robots.txt rule that disallowed one or more tracked crawler or product-control tokens without naming each token explicitly."
¶ cite this findingThe free AI Crawler Checker reports effective rules for documented crawler and product tokens and produces a policy suggestion for human review.
Check Your robots.txt →Chart 3: Adoption by rank decile — the counterintuitive curve
The conventional assumption is that the most popular sites are the most technically sophisticated, and therefore the first to adopt emerging standards. For llms.txt, this is wrong.
The practical implication: llms.txt is currently a deliberate choice, not a default. Mid-tier sites with engaged engineering teams are adopting it faster than the media conglomerates and legacy institutions that dominate the top 1,000. The crawl did not test whether any provider consumes or prefers llms.txt; Google Search explicitly ignores it.
"llms.txt adoption peaks at Tranco ranks 2,001–3,000 (9.1%), not among the highest-ranked 1,000 sites. In this crawl frame, publication of the optional proposal was more common in that mid-tier decile than among the largest legacy publishers."
¶ cite this findingChart 4: Signal adoption rates across the web
AI retrieval systems layer multiple signals. We measured five: HTTPS redirect (the baseline), canonical tag, sitemap declaration in robots.txt, Organization schema, and llms.txt. Each sits at a different maturity point on the adoption curve.
"Only 18.9% of reachable sites in CiteFuel's June 2026 sample published Organization schema. That is an adoption measurement; the study did not test whether the markup changed ranking, recommendation, or citation."
¶ cite this findingChart 5: llms.txt presence was associated with a 20% lower median TTFB
Across all 7,543 reachable sites, we recorded the Time to First Byte (TTFB) of the homepage response. The distribution splits sharply by llms.txt status.
"In this sample, sites with an llms.txt file had a 20% lower median homepage TTFB (203ms vs. 255ms). This correlation does not show that llms.txt caused the difference and may reflect other characteristics of the sites or teams."
¶ cite this findingConfiguration index by industry vertical
Industry classification uses TLD heuristics applied at crawl time (same as
crawl-study/analysis.py). Only verticals with
n ≥ 200 domains are published to ensure statistical credibility.
Two specific verticals clear the bar; the general-web bucket (n = 8,837) serves as the
baseline for all-industry comparisons.
Nonprofit / NGO sites had the highest mean configuration index among published verticals (2.38/5) — edging out Tech / SaaS (mean 2.22/5) despite typically having smaller engineering teams. Both verticals match the general-web median of 2, but nonprofits have a higher share of domains scoring ≥ 3 (39% vs. 34% for tech).
| Vertical | n | Mean score | Median | 25th pct (p25) | 75th pct (p75) | Score distribution (0–5) |
|---|---|---|---|---|---|---|
| All industries (general web) | 8,837 | 2.22 | 2 | 1 | 3 | |
| Nonprofit / NGO .org · .ngo | 395 | 2.38 | 2 | 2 | 3 | |
| Tech / SaaS .io · .dev · .ai · .app | 334 | 2.22 | 2 | 1 | 3 |
Verticals with n < 200 (education n=164, government n=172, media n=93, e-commerce n=5)
are omitted — insufficient sample for reliable inference. Classification: TLD heuristics
(gov/mil → government; edu/ac → education; org/ngo → nonprofit; io/dev/app/ai/tech →
tech; news/media/tv → media; shop/store → e-commerce; all others → general).
Reproducible: python3 crawl-study/benchmark_verticals.py.
"Nonprofit and NGO sites had the highest mean five-signal configuration index: 2.38/5, with 39% scoring 3 or above. Tech / SaaS sites — despite theoretically having stronger engineering resources — score on par with the general web at 2.22/5."
¶ cite this findingThe downloadable study data lets researchers reproduce the five-signal comparisons shown here. The current CiteFuel checker uses a broader methodology and reports its own evidence coverage; its headline is not the same metric as this historical configuration index.
Methodology
Full transparency on what we measured, how we measured it, and where the data falls short — because credibility depends on it.
- Sample frame
- Tranco ranks 1,001–11,000 (10,000 domains), June 2026 list. Ranks 1–1,000 excluded to avoid sampling bias from mega-platforms with outlier infrastructure.
- Crawl date
- June 12, 2026. Point-in-time snapshot. All rates reflect conditions on that date. Re-runs are planned quarterly.
- Crawler
CiteFuelBot/1.0. Full robots.txt compliance. Fetched homepage only (root path). Not a training crawler.- Reachability
- 7,543 / 10,000 homepages returned parseable HTML. Remainder: timeouts, connection errors, non-HTML roots. All 10,000 are in the denominator for blocking-rate calculations.
- Robots.txt opt-outs
- 217 domains disallowed CiteFuelBot. We honored the opt-out and did not fetch their homepage. These are recorded as
robots_excludedin the raw data. - Configuration index
- 5 binary signals: (1) llms.txt present, (2) no disallow rule for the five tracked crawler/product tokens, (3) Organization schema present, (4) canonical tag present, (5) HTTPS redirect active. Each is 1 point. This is not a ranking or visibility score.
- Blocking detection
- Full effective-rule semantics including wildcard fallback (
User-agent: *). Rates computed over all 10,000 domains, not just reachable ones. - TTFB
- Median homepage response time recorded during the same crawl. Computed over reachable sites only. Network conditions affect absolute values; the relative gap is the signal.
Limitations
- Homepage-only. We crawled one URL per domain. Inner pages and subdomains are not captured, so the configuration index cannot characterize the whole site.
- Point-in-time. All rates are a single-date snapshot. The web changes daily. We publish the crawl date prominently and plan quarterly re-runs.
- No WAF probing. We measured effective robots.txt disallow rules, not whether requests from verified crawler networks would be accepted at the CDN/WAF layer. WAF behavior is outside this dataset and no total enforcement rate can be inferred from it.
- Industry breakdown is directional. Industry classification uses TLD heuristics (Tranco carries no category data). Sub-bucket rates (tech, edu, gov) are directional, not authoritative.
- Tranco list dynamics. Tranco is a popularity aggregate, not a curated quality list. Very new domains and CDN edge nodes may appear at these ranks.
Data and reproducibility
The aggregate JSON (stats-2026-06-12.json) is published alongside this report. The full per-domain CSV (10,000 rows) is available on request at research@citefuel.com — we share it freely with journalists, researchers, and developers. The crawl harness code is likewise available on request at the same address.
Publication standards, authorship and review responsibilities, AI-assistance disclosure, funding controls, and conflict handling are documented in CiteFuel's editorial policy. Questions about a figure or interpretation can be submitted through the corrections process.
What this means for site owners
This study is a configuration inventory, not a ranking experiment. Site owners can use it to review intentional robots policies, canonical and HTTPS basics, and whether structured data truthfully matches visible content. Those checks can remove technical ambiguity, but the study provides no evidence that adopting llms.txt or schema will cause an answer engine to mention or cite a site.
For Google specifically, normal indexing and snippet eligibility remain the foundation for AI features; no special AI schema is required, and Google says it ignores llms.txt. Any robots.txt or structured-data change should reflect the site's actual policy and visible content, then be tested after deployment.
Audit your current configuration and sampled visibility with explicit evidence coverage. The current CiteFuel score is a separate internal metric, not a percentile from this five-signal study or a ranking guarantee.
Run Your Free Audit →How to cite this study
Plain text
CiteFuel. "The State of AI Crawlability 2026: We Crawled 10,000 Popularity-Ranked Domains." CiteFuel Research, June 12, 2026. https://citefuel.com/research/state-of-ai-crawlability-2026
BibTeX
@techreport{citefuel2026crawlability, author = {CiteFuel}, title = {The State of AI Crawlability 2026: We Crawled 10,000 Popularity-Ranked Domains}, institution = {CiteFuel}, year = {2026}, month = {jun}, type = {Research Report}, url = {https://citefuel.com/research/state-of-ai-crawlability-2026}, note = {Data: https://citefuel.com/research/data/stats-2026-06-12.json} }
Journalists and researchers: feel free to cite any finding from this page. For the full per-domain dataset, email research@citefuel.com. Browse all CiteFuel Research or review our editorial policy.
Measure your site’s AI-readiness gaps, then review the evidence.
Free 26-check audit. No card. No login. Just a URL — results in ~90 seconds.
Audit my site free →