Claim Verification: "AI hallucinations occur on fewer than 5% of factual questions" — Disproved
收藏资源简介:
Automated fact-verification of the claim: "AI hallucinations occur on fewer than 5% of factual questions" Verdict: DISPROVED Key Findings OpenAI's o3 model hallucinated 33% of the time on the PersonQA benchmark (B1) — nearly 7x the claimed ceiling of 5%. ChatGPT generates hallucinated content in approximately 19.5% of its responses across general testing (B2) — nearly 4x the claimed ceiling. On the AA-Omniscience benchmark (6,000 factual questions across 42 topics), even the best-performing model hallucinates 22% of the time (B3). No major AI model achieves < 5% hallucination on open-ended factual question benchmarks. Sub-5% rates exist only on narrow grounded summarization tasks, not factual QA. Files proof.py — Re-runnable Python verification script proof.md — Structured proof report proof_audit.md — Full verification audit trail proof_narrative.md — Plain-language summary proof.json — Machine-readable structured data Generated by Proof Engine v1.1.0.



