The 2026 State of Generative Engine Optimization
收藏资源简介:
Version 3.0 of The 2026 State of Generative Engine Optimization. Cross-engine divergence: do four production AI search products return the same vendors for the same B2B software buyer question, and is any difference bigger than each engine's own run-to-run noise? 280 buyer questions across 40 categories were put to ChatGPT, Claude, Perplexity and Google AI Overviews on 4 August 2026, and a 25-question stratified subsample was re-run three times on the same engines so between-engine difference could be read against within-engine noise measured on the same questions in the same window. Headline result, reported per engine Engine Repeat pairs Agrees with itself Agrees with other engines Gap (95% CI) Google AI Overviews 62 0.499 0.240 +0.258 (0.178 to 0.338) Claude 74 0.442 0.277 +0.165 (0.082 to 0.242) The length control changes the answer. Engines name very different numbers of vendors per answer (Google AI Overviews 4.8, Claude 8.8, Perplexity 10.7), and an engine compared with itself is length-matched by construction while two different engines are not. Rerunning the same comparison with each answer truncated to its first five named vendors, and again with an overlap coefficient normalised by the smaller set: Google AI Overviews: gap holds under both (+0.167 and +0.193). Claude: gap collapses to +0.075 and +0.002, both intervals crossing zero. On Claude the apparent divergence is not distinguishable from a list-length artefact. Other findings Mean pairwise Jaccard across all engine pairs 0.313 (95% CI 0.295 to 0.330), 590 pair observations. Across 142 questions three engines all answered, 63.3% of distinct vendor mentions came from exactly one engine and 17.5% from all three. Google returned an AI Overview for 92.1% of these buyer questions. 14 of 32 companies testable on more than one engine were visible on some and invisible on others. Changes since v2.0 New research question. v1.0 held one engine constant across 860 answers; v2.0 analysed absence within that single-engine frame. v3.0 is the first version to log which engine produced which answer and to compare engines against each other. New dataset. 853 scored answers, one row per engine-answer, with run id, timestamp, category, question, engine, exact model string, repeat index, status, full answer text, cited URLs, latency and SerpApi calls consumed. Released in full. A within-engine repeat baseline. 25 questions run three times per engine, so divergence can be read against nondeterminism rather than reported in isolation. Per-engine reporting throughout. Pooled figures are demoted, because a pooled within-engine mean is dominated by whichever engine contributed the most repeat pairs. Length controls. Top-k truncation and an overlap coefficient alongside Jaccard. These changed the conclusion on one of the two engines. Two-layer vendor extraction. Deterministic matching using the v1.0 rules, authoritative for universe companies, plus an LLM extractor at temperature 0 that may only add out-of-universe names. Extraction prompt released. An independent verification gate. A separately written script re-derives every headline figure from the raw data and diffs it against the paper text; it must pass before publication. Honest collection reporting. Three of four API accounts ran out of credit mid-collection; one was topped up and finished. Per-engine sample sizes differ and every statistic states the engines and n it was computed on. The four-engine intersection is 18 questions and never carries a headline. Prior-version datasets are included again so this record remains self-contained. Not peer reviewed. Broadcastwell ran this measurement and sells services in the category it measures. It is excluded from the measured sample and from every ranking. Prior versions: v1.0 doi 10.5281/zenodo.21537014, v2.0 doi 10.5281/zenodo.21586091.



