How CiteFuel Scores AI Visibility — Full Methodology
By Hunter Spence, founder — last updated July 12, 2026
Why we publish this: the GEO/AEO space has a snake-oil problem. Tools publish opaque "AI scores" with no explanation of inputs, weights, or validation. We publish ours because (a) transparency builds trust, (b) you should understand what you're optimizing for, and (c) if our weights are wrong, we want to hear about it. Email support@citefuel.com with disagreements. Product scoring changes are recorded on this page. Separate, dataset-based publications live in CiteFuel Research and follow our editorial policy and corrections process.
Score 2.0 (current engine): your published GEO score is a blend of two layers —
Readiness (R), the deterministic on-site check table below
(26 checks worth 129 points,
R = earned ÷ available × 100, same partial/skip rules as before), and
Visibility (V), a live-sampled presence rate: we generate a disclosed synthetic prompt
set about your category, sample an AI-answer engine across N runs, and report the
share that produced a positive brand-presence reading — as a rate with a 90% bootstrap confidence interval.
A text adapter generally counts a brand appearance; a supported source/reference match may also count.
This is not a uniform recommendation or citation rate. Individual outputs vary, so repeated fixed-prompt
sampling and an interval quantify, but do not remove, uncertainty. Your published
score is a confidence-weighted blend of R and V — nominally 0.65×R + 0.35×V, but V's
actual weight scales with how many samples it's based on (see "Confidence-weighted blend" below for the
full formula and why). If Visibility wasn't sampled this run (rate-limited, errored, or genuinely not
yet run), the score is R alone, and the report says so explicitly — never a guessed or assumed V.
Google-specific boundary: Google says its AI features use the same foundational Search requirements, require normal indexing and snippet eligibility, need no special AI schema, and ignore llms.txt. See Google's AI features guidance and generative AI optimization guide. CiteFuel's checks and scores are internal measurements, not Google ranking signals.
Hard gates: two implemented failures are severe enough that no amount of Readiness or Visibility credit offsets them: the HTTPS/TLS check fails, or an AI assistant asked cold about the brand states a claim that clearly contradicts the site's own evidence. Either condition caps the published score at 39 (grade D), regardless of what R and V would otherwise blend to. The WAF row is a CiteFuel-origin user-agent response comparison, not verified crawler-network evidence; it remains coverage-critical and can lower Readiness, but it is not a 39/D hard-cap condition.
The 6-engine visibility matrix
Visibility (V) attempts up to 6 AI-answer engines. Authentication, billing, rate limits, and provider health can change which lanes complete — this configuration snapshot is not a promise of readings:
| Engine | Status |
|---|---|
| ChatGPT | Configured lane — only completed readings count |
| Claude (web search) | Configured lane — authentication/provider failures remain unavailable |
| Gemini | Configured lane — only completed readings count |
| Perplexity | Configured lane — only completed readings count |
| Grok | Configured lane — only completed readings count |
| Google AI Overviews | Disabled by default pending verified DataForSEO billing health; disabled or unavailable lanes produce no evidence |
A report's sample count only counts completed readings. An unavailable engine is neither a zero nor a pass; it is reported as unavailable and excluded from the measured denominator. The report's own engine, timestamp, locale, N, confidence interval, and provider status are authoritative for that run.
Confidence-weighted blend — why your V weight isn't always 35%
The nominal split is 65% Readiness / 35% Visibility, but Visibility's actual weight in your headline scales with how many samples it's based on — a single unlucky or lucky sampled answer can swing a pooled presence rate by a large amount. Repeated answers can vary even for a fixed prompt, so a 10-sample free-tier read shouldn't count for as much as a robust 48-sample paid-tier read.
confidence = min(1, n_runs ÷ 20) — N ≥ 20 is the full-confidence weighting
threshold, where Visibility reaches its full nominal 35% weight. Below that, V's effective weight is 0.35 × confidence and R
picks up the difference, so the two always sum to 100%. At 0 samples this degenerates to R alone (100%
Readiness, 0% Visibility) — the same behavior as before Visibility sampling existed at all. At 20+
samples it's byte-identical to the nominal 65/35 split.
Visibility here means positive brand-presence readings, not a uniform citation or recommendation rate.
Confidence-weighting prevents a small sample from taking the full nominal 35% of the formula. For
example, Readiness 100 and Visibility 0 at pooled N=10 yields
0.825 × 100 + 0.175 × 0 = 82.5. That low-N output is provisional: it is not
evidence of a grade or certification and is not comparable with an N ≥ 20 full-confidence
audit. Reports disclose the effective weight and N; the separate evidence gate also prevents low-N
weighting from producing a new S/90+ certification.
Coverage and S/90+ certification gate
The displayed formula score is not enough to earn an S grade. New reports can publish 90+ only when the
unrounded formula is at least 90.0 and all of these are true: at least four visibility
engines produced readings; pooled N ≥ 20 (the full-confidence weighting threshold); at least 90%
of Readiness check weight was actually measured; and no critical crawl/index/citation measurement was
skipped. Visibility must also be fresh: a cached reading cannot earn a new S/90+ certification. A robust
badge additionally requires N ≥ 40. If display rounding produces 90 while the raw formula is
below 90.0, or if an evidence gate fails, the report shows the raw unrounded value, display-rounded value,
and gate reasons, then caps the published score at 89/A. Missing evidence never becomes successful evidence.
Every report lists measured weight, skipped checks, healthy engines, attempted engines, unavailable
provider reasons, pooled N, confidence interval, and the exact gate reasons. Cached results disclose the
original Visibility measurement time and cache age. Authority coverage is
reported separately in the SEO track; an OpenPageRank not_found response means authority
is unknown, not zero and not passed.
Mention rank & cited-source rank
Alongside the presence rate, each engine's reading includes two rank signals computed from the same sampled answers (no extra API calls):
- Mention rank — your brand's position among every brand name mentioned in the engine's answer, ordered by which name appears earliest in the text. Rank 1 means you were the first brand named; a higher number means competitors were named before you. No rank at all means your brand never came up in that answer.
- Cited-source rank — your domain's position within the engine's own list of cited sources/references for that answer (when the engine surfaces one), same ordering convention. This is a different signal from mention rank: an engine can name your brand without ever citing your site as a source, or vice versa.
Your report shows the median of each rank across every sampled run for that engine — a single run's rank is noisy, so the median across N runs is less sensitive to one outlier. It still carries sampling uncertainty and should be compared across identically configured audits.
Every check below still keeps the same partial/skip rules as before: a partial result earns half the check's weight, and a skipped check (e.g. AI sampling temporarily unavailable, or a Wikipedia article that doesn't exist for your brand) is excluded from the denominator. Coverage and skips are disclosed separately so a smaller denominator is never mistaken for stronger evidence.
Category 1 — AI Crawler Access · 6 checks · 28 pts (21.7%)
The access layer tests whether each named crawler or product token is allowed by robots.txt and whether a request using documented crawler identification receives usable content. A blocked retrieval crawler cannot fetch the tested URL directly, but that result alone does not prove a brand can never appear through other indexes or sources.
ai_robots_gptai_robots_claudeai_robots_geminiai_robots_perplexityai_robots_metawaf_block_heuristicCategory 2 — llms.txt Presence & Quality · 2 checks · 4 pts (3.1%)
llms.txt is a voluntary proposal for publishing a curated Markdown site map. We check optional-file syntax and link hygiene for services that choose to consume it. Google Search explicitly ignores llms.txt, so this check is not a Google ranking, AI Overview, or AI Mode lever.
llms_txt_presentllms_txt_qualityCategory 3 — Schema Markup AI-Readiness · 3 checks · 15 pts (11.6%)
We validate whether JSON-LD parses, uses a schema.org context, has safe absolute identity-reference URLs, and avoids obvious injection strings. This automated check does not prove that every claim matches visible content; publishers remain responsible for that. Structured data can support search understanding and eligibility for supported rich-result features, but Google requires no special AI schema and valid markup does not guarantee ranking or citation.
jsonld_presentjsonld_coveragejsonld_sanityCategory 4 — Passage Citability & Entity · 6 checks · 41 pts (31.8%)
Passage-level citability is CiteFuel's internal heuristic for clarity, evidence, and standalone readability; it is not an engine ranking factor or citation probability. Entity footprint estimates how distinct the brand appears from measured public evidence.
citability_passageentity_footprintanswer_first_structurequotation_stat_densityoffsite_mentionswikipedia_qualityCategory 5 — Technical Foundation · 9 checks · 41 pts (31.8%)
The classical technical foundation shared with search: canonical correctness, HTTPS integrity, sitemap discoverability, Open Graph completeness, mobile viewport, and PageSpeed metrics. PageSpeed field data is preferred; Lighthouse lab fallback is lab evidence, not real-user Core Web Vitals proof.
freshness_signalscanonical_tagcwv_lcpcwv_clscwv_inphttps_tlssitemap_presentog_tagsviewport_metaEngine v1 (legacy — reports scored before 2026-07-03)
Reports generated before the Score 2.0 cutover (2026-07-03) were scored against the table below —
23 checks worth 111 points, a single
Readiness-only score with no Visibility blend or hard gates. Those reports still render with this
table's weights for an honest like-for-like comparison; every audit run today shows its v1 reading
too (report.score_v1) alongside the current Score 2.0 headline.
v1 Category 1 — AI Crawler Access · 6 checks · 28 pts (25.2%)
The access layer tests whether each named crawler or product token is allowed by robots.txt and whether a request using documented crawler identification receives usable content. A blocked retrieval crawler cannot fetch the tested URL directly, but that result alone does not prove a brand can never appear through other indexes or sources.
ai_robots_gptai_robots_claudeai_robots_geminiai_robots_perplexityai_robots_metawaf_block_heuristicv1 Category 2 — llms.txt Presence & Quality · 2 checks · 13 pts (11.7%)
llms.txt is a voluntary proposal for publishing a curated Markdown site map. We check optional-file syntax and link hygiene for services that choose to consume it. Google Search explicitly ignores llms.txt, so this check is not a Google ranking, AI Overview, or AI Mode lever.
llms_txt_presentllms_txt_qualityv1 Category 3 — Schema Markup AI-Readiness · 3 checks · 15 pts (13.5%)
We validate whether JSON-LD parses, uses a schema.org context, has safe absolute identity-reference URLs, and avoids obvious injection strings. This automated check does not prove that every claim matches visible content; publishers remain responsible for that. Structured data can support search understanding and eligibility for supported rich-result features, but Google requires no special AI schema and valid markup does not guarantee ranking or citation.
jsonld_presentjsonld_coveragejsonld_sanityv1 Category 4 — Passage Citability & Entity · 2 checks · 13 pts (11.7%)
Passage-level citability is CiteFuel's internal heuristic for clarity, evidence, and standalone readability; it is not an engine ranking factor or citation probability. Entity footprint estimates how distinct the brand appears from measured public evidence.
citability_passageentity_footprintv1 Category 5 — Technical Foundation · 8 checks · 28 pts (25.2%)
The classical technical foundation shared with search: canonical correctness, HTTPS integrity, sitemap discoverability, Open Graph completeness, mobile viewport, and PageSpeed metrics. PageSpeed field data is preferred; Lighthouse lab fallback is lab evidence, not real-user Core Web Vitals proof.
canonical_tagcwv_lcpcwv_clscwv_inphttps_tlssitemap_presentog_tagsviewport_metav1 Category 6 — Live AI Answer Presence · 2 checks · 14 pts (12.6%)
The outcome layer samples live answer engines and records whether the brand produced a positive presence reading under each adapter. Presence is not uniformly a recommendation or citation. Unavailable readings are excluded and disclosed rather than treated as zeroes or passes.
ai_brand_sample_chatgptai_brand_sample_perplexitySEO Track B — a separate, thin sibling score
Alongside the GEO score above, every audit also computes a lightweight SEO score —
classic on-page and technical search-visibility signals, worth 95 points
across 12 checks. It's reported separately (report.seo) and
never mixed into your GEO score or grade — the two are scored independently. Checks marked
"reuses your GEO result" read an existing GEO check's verdict rather than re-running it.
SEO 1 — On-Page SEO · 5 checks · 41 pts (43.2%)
title_tagmeta_descriptionh1_structureindexabilityredirect_healthSEO 2 — Technical SEO (shared signals — reused from your GEO check results) · 5 checks · 34 pts (35.8%)
canonicalcanonical_tag check result — not re-run. sitemapsitemap_present check result — not re-run. https_mixedhttps_tls check result — not re-run. viewportviewport_meta check result — not re-run. cwv_lcpcwv_lcp check result — not re-run. SEO 3 — Domain Authority · 2 checks · 20 pts (21.1%)
domain_authoritybrand_serp_ownership
A 13th check, Referring domains (backlinks)
(backlink_referring_domains, 8 pts), runs on paid audits and
enters their measured SEO denominator. It is not included in the free 95-pt table.
App AI-Visibility — a separate scale, for iOS App Store links
Paste an App Store link instead of a domain and CiteFuel runs a different, purpose-built
14-row audit registry, never mixed into the website GEO or SEO scores above. The current
app_v2 headline normalizes 10 scored Readiness rows worth
75 available points, then blends that Readiness result with the
separately sampled Visibility layer. A partial earns half weight and a skipped scored row is excluded from
the available denominator. Two developer-site rows — optional app schema and the llms.txt/FAQ/answer-first
inventory — are reported but informational and earn zero Readiness credit. The two live sampling rows,
ai_recommendation_sampling and competitor_sov, describe one shared Visibility sample
and also earn no fixed Readiness points. The historical 100-point flat checklist remains in
report JSON only for explicitly labeled legacy comparison.
App 1 — App Store Listing · 4 rows · 37 legacy checklist pts
What the approved App Store listing itself says: whether the subtitle uses a descriptive category term, the description opens clearly, and relevant listing language is covered without stuffing. CiteFuel does not claim this proves query demand.
listing_subtitle_query_matchlisting_description_answer_firstlisting_keyword_coveragescreenshot_captionsApp 2 — Ratings Signal · 2 rows · 18 legacy checklist pts
Rating volume and an age-adjusted lifetime-average ratings-per-month proxy, benchmarked against a live category corpus. The latter does not measure recent rating velocity or prove current momentum.
rating_volumerating_lifetime_average_rateApp 3 — Freshness · 1 rows · 8 legacy checklist pts
Time since the last shipped update. An app that looks abandoned reads as abandoned to both users and AI answer engines regardless of how good the listing copy is.
update_freshnessApp 4 — Developer Site · 3 rows · 19 legacy checklist pts
If a developer website is listed: does it resolve, link clearly to the App Store listing, expose important content in visible text, and use truthful applicable structured data. Optional files and schema are never treated as ranking guarantees.
dev_website_presentdev_website_app_schemadev_website_geo_subsetApp 5 — Off-Site Presence · 4 rows · 18 legacy checklist pts
Whether the app is mentioned outside the App Store — Reddit, "best apps" listicles, and live AI-answer sampling with competitor share-of-voice. Each row scores only when its source returns usable evidence; unavailable sources are disclosed as skips, not zeroes or passes.
offsite_reddit_mentionsoffsite_listicle_presenceai_recommendation_samplingcompetitor_sovGrades & scoring tiers
| Score | Grade | What it means |
|---|---|---|
| 90-100 | S | High measured CiteFuel score with the S/90+ coverage gate satisfied; not a ranking guarantee. |
| 75-89 | A | Strong measured score or an otherwise-high score awaiting sufficient evidence coverage. |
| 60-74 | B | Meaningful gaps. P1 fixes recommended within 30 days. |
| 45-59 | C | Significant gaps in CiteFuel's measured readiness rubric. |
| 0-44 | D | Critical tested failures or many readiness gaps; not a forecast of ranking or citation. |
Severity tiers
Every failing check is assigned an operational priority: P0 is a tested access, security, or indexability failure for the named surface; P1 is an important internal-readiness gap; and P2 is a refinement. These labels prioritize remediation. They are not estimates of ranking or citation probability. Free reports show the full priority-ranked gap list; paid audits deliver the fix files.
Changelog
- 2026-07-12 — added S/90+ evidence-coverage gates (4 healthy engines, N≥20, ≥90% measured check weight, no critical skip; N≥40 for robust badge); paid SEO now activates its paid authority row; provider failures and OpenPageRank not-found results remain explicit unknowns.
- 2026-07-06 — reweighted
sitemap_present(4→10 pts) andcanonical_tag(4→9 pts) in the Readiness table; both signals were underweighted relative to their impact on crawl discoverability. Readiness table now totals 129 pts, up from 118. Check count unchanged (26). - v2.0 (2026-07-03) — Score 2.0: published score became 0.65×Readiness + 0.35×Visibility (26 checks / 118 pts Readiness table, up from 23/111 — added freshness signals, answer-first structure, quotation/stat density, off-site mentions, and Wikipedia quality; moved live AI-presence sampling out of Readiness and into the new Visibility layer), with a 39/D cap for the implemented HTTPS/TLS and evidence-contradiction hard gates. v1 stays computed and shown for comparison (
report.score_v1). - v1.0 (2026-06-11) — initial 23-check framework. We update when major AI crawlers document behavior changes; weight changes are recorded here.