Clear Cited AI Visibility Index: share of model across five AI engines, nine B2B software categories
收藏资源简介:
An open, pre-registered measurement of which products generative AI engines name and cite when buyers ask for software recommendations. Nine B2B software categories were measured on 2026-07-01 across five AI engines — ChatGPT (OpenAI), Perplexity, Claude (Anthropic), Gemini (Google) and Grok (xAI) — using a fixed protocol of 10 prompts × 10 runs per engine, yielding 500 AI answers per category. The dataset covers 4,497 AI answers across 90 ranked products. It also carries 34,801 extracted citation records as captured — including, we subsequently found, a substantial proportion of platform redirect wrappers and near-zero recovery from two of the five engines. The citation layer is published exactly as measured so that the defect is independently checkable; no domain-level conclusion should be drawn from it. For each product we report share of model (the proportion of qualifying AI answers naming the brand), a 95% Wilson confidence interval on that share, appearance rate, citation rate, raw mention count and answer total. The accompanying JSON additionally carries per-engine breakdowns, full citation-source records, capture timestamps and model identifiers. Prompts and the brand universe were hashed and frozen before the run; the pre-registration hashes are included in each JSON record. Products cannot pay to appear or to rank. Google AI Overviews and Microsoft Copilot are treated as answer surfaces and measured separately. They are never summed into the five-engine share of model. **Categories:** AI observability tools · API platforms · CI/CD platforms · CRM software · databases · feature flag platforms · incident management platforms · product analytics platforms · vector databases. **Limitations.** A single measurement run captures one point in time; model versions change beneath the measurement, which is why every record is dated at capture. Ten prompts per category is a deliberate trade between coverage and cost, and does not span every phrasing a buyer might use. The confidence intervals describe sampling variation within this protocol; they do not account for model drift between runs. **Citation extraction is unreliable in this release.** One engine returned zero distinct cited domains in all nine categories and a second returned one to two against a list cap of five; roughly two in five extracted citations resolve to a platform redirect wrapper rather than a publisher domain. The `citation_rate` column and the `citation_sources` block should not be used for domain-level analysis until a corrected release is published. On the term "citation". At least three distinct things are reported as citations in thisfield: a source retrieved into a model's context versus one cited in the answer (Ahrefs);a source linked versus a brand named in the text (Semrush); and knowledge held in the modelversus retrieved at query time (Evertune). Published estimates of the same nominal quantitydiffer by more than an order of magnitude, which is what one would expect when threedifferent measurements share one word. The citation records in thisrelease are extraction attempts from the returned answer payload - that is, an attempt atthe second sense - and, as noted above, that extraction is currently unreliable. Thisrelease does not measure retrieval into context, which is not observable from an APIresponse.



