CiteFuel Research · Five-signal configuration inventory

The State of AI Crawlability 2026:
We Crawled 10,000 Popularity-Ranked Domains

Only 6.9% of 10,000 popularity-ranked domains (Tranco ranks 1,001–11,000) publish an llms.txt file. 12.5% disallow at least one tracked crawler or product-control token — and 1.9% do so through a wildcard rather than a named rule. This is a point-in-time configuration study; it did not measure rankings, citations, recommendations, or traffic.

By CiteFuel Research Team · Responsible editor: Hunter Spence · Published · Updated · n = 7,543 reachable homepages · Download aggregate data (JSON) · Editorial policy · Report a correction
▸ Key Findings — 8 quotable facts from this study
  • 1 Only 6.9% of 10,000 popularity-ranked domains at Tranco ranks 1,001–11,000 publish an llms.txt file, according to CiteFuel's June 2026 crawl.
  • 2 One in eight reachable websites (12.5%) disallowed at least one tracked crawler or product token in robots.txt. This is a policy measurement, not proof of product-level invisibility.
  • 3 1.9% of reachable sites used a wildcard Disallow: / rule that affected tracked agents without naming each token explicitly.
  • 4 Only 2.9% of reachable sites score 5 out of 5 on CiteFuel's configuration index. The most common score is 2/5, and 47 sites score 0/5; the index does not measure external visibility.
  • 5 Just 18.9% of reachable sites publish Organization schema. Its absence is a configuration observation, not a ranking or citation failure.
  • 6 llms.txt adoption is not highest among the biggest websites: it peaks at ranks 2,001–3,000 (9.1%) and is flat or declining among the very top 1,000 sites. Fast-moving mid-tier sites are the early adopters.
  • 7 Sites with an llms.txt file had a 20% faster median TTFB in this sample (203ms vs. 255ms). This is a correlational association that may reflect confounding factors; it does not show that llms.txt caused faster responses.
  • 8 GPTBot had the highest disallow rate among the tracked tokens (11.3%), followed by ClaudeBot (10.3%), Google-Extended (9.5%), anthropic-ai (6.4%), and PerplexityBot (6.2%). Google-Extended is a non-requesting product-control token, not a crawler user agent.
Section 01 — Overview

The five-signal configuration distribution is wide

In June 2026, CiteFuel crawled 10,000 domains from Tranco ranks 1,001–11,000 — a large popularity-ranked cohort, not a statistically representative sample of the open web — to inventory five crawl and page-configuration signals. Of the 10,000 domains queried, 7,543 returned parseable homepages. The rest timed out, returned non-HTML roots, or disallowed our crawler (217 domains honored; see Methodology).

6.9%
have llms.txt

693 of ~10,000 domains have any AI-guidance file

12.5%
disallow ≥1 tracked token

crawler and product-control tokens combined

2.9%
score 5/5

216 sites publish all five measured items; mode is 2/5

18.9%
Organization schema

4 in 5 sites lack a structured entity anchor

The measured conclusion is narrower: most sites did not publish every item in CiteFuel's five-signal configuration checklist. This crawl did not measure rankings, recommendations, citations, or traffic, so a low checklist score is not proof that a site is invisible to AI.

This historical dataset provides a reproducible configuration benchmark. The current free CiteFuel audit uses a broader, separately documented methodology and should not be treated as this five-item index.

Check Your Site Free →
Section 02 — Configuration Index Distribution

Chart 1: Five-signal configuration index (0–5)

Each reachable site was scored on five binary signals: llms.txt present, no disallow rule for the tracked crawler/product tokens, Organization schema present, canonical tag present, and HTTPS redirect active. Scores range from 0 (none) to 5 (all five).

Configuration Index Distribution — 7,543 reachable sites n = 7543
47
0/5
1,253
1/5
2,664
2/5
2,276
3/5
1,087
4/5
216
5/5
Takeaway: Score 2/5 is the mode — sites with HTTPS and a canonical tag but nothing included in this checklist. The 5/5 tier (216 sites, 2.9%) published all measured configuration items; the 0/5 tail (47 sites) published none. Neither result proves external visibility or invisibility.

"Only 2.9% of reachable sites in CiteFuel's June 2026 sample published all five items in its configuration index. The most common score was 2 out of 5. The index did not measure external search or answer-engine visibility."

¶ cite this finding
Section 03 — Crawler & Product-Token Policy

Chart 2: Per-token robots.txt disallow rates

We checked every domain's robots.txt against five crawler or product-control tokens using effective-rule semantics — including wildcard fallback. A token is counted as disallowed when the effective rule is Disallow: / for that user-agent, whether the token is named explicitly or inherited from a wildcard group. Google-Extended is different from the request-capable crawlers in this chart: Google documents it as a standalone product token that does not make HTTP requests. Its setting does not affect inclusion or ranking in Google Search, including AI Overviews and AI Mode.

Per-Token robots.txt Disallow Rates — % of ~10,000 domains robots.txt effective-rule analysis
GPTBot
11.3%
ClaudeBot
10.3%
anthropic-ai
6.4%
Google-Extended
9.5%
PerplexityBot
6.2%
Takeaway: GPTBot had the highest measured disallow rate (11.3%), followed by ClaudeBot (10.3%). These rates describe robots.txt policy only. They do not establish a causal relationship with ranking, recommendation, citation, or traffic, and Google-Extended is not a crawler.

The silent blocking problem

Across the sample, 1.9 percentage points had a wildcard rule that disallowed tracked tokens without naming those tokens explicitly. For example, GPTBot can inherit User-agent: * and Disallow: / even when the string "GPTBot" does not appear in the file. This is an effective-rule observation, not evidence about the administrator's intent.

"In CiteFuel's June 2026 sample, 1.9% of domains had a wildcard robots.txt rule that disallowed one or more tracked crawler or product-control tokens without naming each token explicitly."

¶ cite this finding

The free AI Crawler Checker reports effective rules for documented crawler and product tokens and produces a policy suggestion for human review.

Check Your robots.txt →
Section 04 — llms.txt Adoption by Rank

Chart 3: Adoption by rank decile — the counterintuitive curve

The conventional assumption is that the most popular sites are the most technically sophisticated, and therefore the first to adopt emerging standards. For llms.txt, this is wrong.

llms.txt Adoption by Rank Decile Tranco ranks 1,001–11,000 · 1,000 domains per bar
avg 6.9%
6.2%
9.1%
7.6%
7.1%
7.9%
6.9%
6.4%
5.8%
5.8%
6.2%
1K–2K
2K–3K
3K–4K
4K–5K
5K–6K
6K–7K
7K–8K
8K–9K
9K–10K
10K–11K
Takeaway: llms.txt peaks at ranks 2,001–3,000 (9.1% — highlighted in amber). The very top 1,000 sites (d1, 6.2%) and the largest sites (d10, 6.2%) show the same adoption rate — consistent with llms.txt being a tech-savvy early-adopter signal, not a size-correlated one.

The practical implication: llms.txt is currently a deliberate choice, not a default. Mid-tier sites with engaged engineering teams are adopting it faster than the media conglomerates and legacy institutions that dominate the top 1,000. The crawl did not test whether any provider consumes or prefers llms.txt; Google Search explicitly ignores it.

"llms.txt adoption peaks at Tranco ranks 2,001–3,000 (9.1%), not among the highest-ranked 1,000 sites. In this crawl frame, publication of the optional proposal was more common in that mid-tier decile than among the largest legacy publishers."

¶ cite this finding
Section 05 — Signal Adoption

Chart 4: Signal adoption rates across the web

AI retrieval systems layer multiple signals. We measured five: HTTPS redirect (the baseline), canonical tag, sitemap declaration in robots.txt, Organization schema, and llms.txt. Each sits at a different maturity point on the adoption curve.

Configuration Signal Adoption — % of reachable sites n = 7543 reachable homepages
HTTPS redirect
86%
Canonical tag
49.5%
Sitemap in robots.txt
38%
Organization schema
18.9%
llms.txt
6.9%
Takeaway: HTTPS (86%) is table stakes — nearly universal. Canonical (49.5%) and sitemaps (38%) are mainstream but not universal. Organization schema (18.9%) and llms.txt (6.9%) are early-adopter territory. The gap between HTTPS and llms.txt adoption is 79 percentage points. That is an adoption gap inside this five-signal checklist, not a visibility opportunity estimate.

"Only 18.9% of reachable sites in CiteFuel's June 2026 sample published Organization schema. That is an adoption measurement; the study did not test whether the markup changed ranking, recommendation, or citation."

¶ cite this finding
Section 06 — Speed Correlation

Chart 5: llms.txt presence was associated with a 20% lower median TTFB

Across all 7,543 reachable sites, we recorded the Time to First Byte (TTFB) of the homepage response. The distribution splits sharply by llms.txt status.

Median TTFB: Sites With vs. Without llms.txt n = 7543 · median homepage response time
203ms
With
llms.txt
255ms
Without
llms.txt
Takeaway: The 52ms median gap is not a causal claim — faster sites likely have engineering teams that also adopt emerging standards early. llms.txt is a marker of technical maturity, not its cause.

"In this sample, sites with an llms.txt file had a 20% lower median homepage TTFB (203ms vs. 255ms). This correlation does not show that llms.txt caused the difference and may reflect other characteristics of the sites or teams."

¶ cite this finding
Section 06b — Vertical Benchmarks

Configuration index by industry vertical

Industry classification uses TLD heuristics applied at crawl time (same as crawl-study/analysis.py). Only verticals with n ≥ 200 domains are published to ensure statistical credibility. Two specific verticals clear the bar; the general-web bucket (n = 8,837) serves as the baseline for all-industry comparisons.

Nonprofit / NGO sites had the highest mean configuration index among published verticals (2.38/5) — edging out Tech / SaaS (mean 2.22/5) despite typically having smaller engineering teams. Both verticals match the general-web median of 2, but nonprofits have a higher share of domains scoring ≥ 3 (39% vs. 34% for tech).

Vertical n Mean score Median 25th pct (p25) 75th pct (p75) Score distribution (0–5)
All industries (general web) 8,837 2.22 2 1 3 1: 35% 2: 26% 3: 24% 4: 13% 5: 2%
Nonprofit / NGO .org · .ngo 395 2.38 2 2 3 1: 19% 2: 41% 3: 27% 4: 12% 5: 2%
Tech / SaaS .io · .dev · .ai · .app 334 2.22 2 1 3 1: 40% 2: 22% 3: 19% 4: 15% 5: 4%

Verticals with n < 200 (education n=164, government n=172, media n=93, e-commerce n=5) are omitted — insufficient sample for reliable inference. Classification: TLD heuristics (gov/mil → government; edu/ac → education; org/ngo → nonprofit; io/dev/app/ai/tech → tech; news/media/tv → media; shop/store → e-commerce; all others → general). Reproducible: python3 crawl-study/benchmark_verticals.py.

"Nonprofit and NGO sites had the highest mean five-signal configuration index: 2.38/5, with 39% scoring 3 or above. Tech / SaaS sites — despite theoretically having stronger engineering resources — score on par with the general web at 2.22/5."

¶ cite this finding

The downloadable study data lets researchers reproduce the five-signal comparisons shown here. The current CiteFuel checker uses a broader methodology and reports its own evidence coverage; its headline is not the same metric as this historical configuration index.

Section 07 — Methodology

Methodology

Full transparency on what we measured, how we measured it, and where the data falls short — because credibility depends on it.

Sample frame
Tranco ranks 1,001–11,000 (10,000 domains), June 2026 list. Ranks 1–1,000 excluded to avoid sampling bias from mega-platforms with outlier infrastructure.
Crawl date
June 12, 2026. Point-in-time snapshot. All rates reflect conditions on that date. Re-runs are planned quarterly.
Crawler
CiteFuelBot/1.0. Full robots.txt compliance. Fetched homepage only (root path). Not a training crawler.
Reachability
7,543 / 10,000 homepages returned parseable HTML. Remainder: timeouts, connection errors, non-HTML roots. All 10,000 are in the denominator for blocking-rate calculations.
Robots.txt opt-outs
217 domains disallowed CiteFuelBot. We honored the opt-out and did not fetch their homepage. These are recorded as robots_excluded in the raw data.
Configuration index
5 binary signals: (1) llms.txt present, (2) no disallow rule for the five tracked crawler/product tokens, (3) Organization schema present, (4) canonical tag present, (5) HTTPS redirect active. Each is 1 point. This is not a ranking or visibility score.
Blocking detection
Full effective-rule semantics including wildcard fallback (User-agent: *). Rates computed over all 10,000 domains, not just reachable ones.
TTFB
Median homepage response time recorded during the same crawl. Computed over reachable sites only. Network conditions affect absolute values; the relative gap is the signal.

Limitations

  • Homepage-only. We crawled one URL per domain. Inner pages and subdomains are not captured, so the configuration index cannot characterize the whole site.
  • Point-in-time. All rates are a single-date snapshot. The web changes daily. We publish the crawl date prominently and plan quarterly re-runs.
  • No WAF probing. We measured effective robots.txt disallow rules, not whether requests from verified crawler networks would be accepted at the CDN/WAF layer. WAF behavior is outside this dataset and no total enforcement rate can be inferred from it.
  • Industry breakdown is directional. Industry classification uses TLD heuristics (Tranco carries no category data). Sub-bucket rates (tech, edu, gov) are directional, not authoritative.
  • Tranco list dynamics. Tranco is a popularity aggregate, not a curated quality list. Very new domains and CDN edge nodes may appear at these ranks.

Data and reproducibility

The aggregate JSON (stats-2026-06-12.json) is published alongside this report. The full per-domain CSV (10,000 rows) is available on request at research@citefuel.com — we share it freely with journalists, researchers, and developers. The crawl harness code is likewise available on request at the same address.

Publication standards, authorship and review responsibilities, AI-assistance disclosure, funding controls, and conflict handling are documented in CiteFuel's editorial policy. Questions about a figure or interpretation can be submitted through the corrections process.

Section 08 — Implications

What this means for site owners

This study is a configuration inventory, not a ranking experiment. Site owners can use it to review intentional robots policies, canonical and HTTPS basics, and whether structured data truthfully matches visible content. Those checks can remove technical ambiguity, but the study provides no evidence that adopting llms.txt or schema will cause an answer engine to mention or cite a site.

For Google specifically, normal indexing and snippet eligibility remain the foundation for AI features; no special AI schema is required, and Google says it ignores llms.txt. Any robots.txt or structured-data change should reflect the site's actual policy and visible content, then be tested after deployment.

Audit your current configuration and sampled visibility with explicit evidence coverage. The current CiteFuel score is a separate internal metric, not a percentile from this five-signal study or a ranking guarantee.

Run Your Free Audit →
Section 09 — Citation

How to cite this study

Plain text

CiteFuel. "The State of AI Crawlability 2026: We Crawled 10,000 Popularity-Ranked Domains."
CiteFuel Research, June 12, 2026. https://citefuel.com/research/state-of-ai-crawlability-2026

BibTeX

@techreport{citefuel2026crawlability,
  author       = {CiteFuel},
title        = {The State of AI Crawlability 2026: We Crawled 10,000 Popularity-Ranked Domains},
  institution  = {CiteFuel},
  year         = {2026},
  month        = {jun},
  type         = {Research Report},
  url          = {https://citefuel.com/research/state-of-ai-crawlability-2026},
  note         = {Data: https://citefuel.com/research/data/stats-2026-06-12.json}
}

Journalists and researchers: feel free to cite any finding from this page. For the full per-domain dataset, email research@citefuel.com. Browse all CiteFuel Research or review our editorial policy.

Measure your site’s AI-readiness gaps, then review the evidence.

Free 26-check audit. No card. No login. Just a URL — results in ~90 seconds.

Audit my site free →