遇见数据集

The State of AI Answers: a longitudinal archive of AI-engine answer sets

收藏
Zenodo2026-07-08 更新2026-08-02 收录
官方服务:

资源简介:

A longitudinal archive of 595 AI-engine answers to 42 buyer-intent prompts across 4 product categories and 3 engines, with repeated runs per prompt to expose run-to-run variance. Each row is one complete engine answer: the full answer text, the tracked brands it named (with first-mention order), and the web domains the engine cited as grounding. Prompts were run multiple times per engine per pull, so the same question appears with diverging answers. That divergence is a first-class part of the data, and the reason single-run "AI visibility checks" mislead. Engines were queried through API and proxy endpoints (DataForSEO), not the consumer apps, with a fixed prompt list per basket and typically 3 repeat runs per prompt per engine per pull. - `google_aio`: Google AI Overviews via the DataForSEO SERP endpoint. Where recorded, conditions were geo US, language en, desktop. - `chatgpt_d4s`: ChatGPT via the DataForSEO AI Optimization LLM Responses endpoint with web grounding, model `gpt-5.4-mini`. - `perplexity_d4s`: Perplexity via the same endpoint, model `sonar`. Citations come from the endpoints' source annotations (real returned source URLs, not inferred from text) and are normalized to registrable domains. The archive is append-only: rows are never edited or overwritten after collection, and basket composition changes are logged append-only (see limitations).

提供机构:
Zenodo
创建时间:
2026-07-08
二维码
社区交流群
二维码
科研交流群
商业服务