遇见数据集

Claim Verification: "Current AI systems in 2026 have near-zero hallucinations and human-level reasoning across most domains." — Disproved

收藏
Zenodo2026-04-09 更新2026-06-05 收录
官方服务:

资源简介:

Automated fact-verification of the claim: "Current AI systems in 2026 have near-zero hallucinations and human-level reasoning across most domains." Verdict: DISPROVED Key Findings AI hallucination rates are far from near-zero: The Vectara Hallucination Leaderboard (2025) shows Gemini-3-pro at 13.6% and all major frontier models (Claude Sonnet 4.5, GPT-5, GPT-OSS-120B, Grok-4, DeepSeek-R1) above 10% — more than an order of magnitude above any "near-zero" threshold. AI falls dramatically short of human-level reasoning on rigorous benchmarks: ARC-AGI-3 (launched March 26, 2026) shows every tested frontier model scoring below 1% against a human baseline of 100%; the best commercial AI scored 0.37%. Expert-domain accuracy exposes the gap: On Humanity's Last Exam (a 2,500-question test spanning 100+ academic disciplines), GPT-4o achieved 2.7% and o1 achieved only 8% — far below expert human performance. Both sub-claims are independently disproved by 3 verified sources each: 6 of 6 citations verified live with full-quote matching; no counter-evidence found that overturns either finding. Files proof.py — Re-runnable Python verification script proof.md — Structured proof report proof_audit.md — Full verification audit trail proof_narrative.md — Plain-language summary proof.json — Machine-readable structured data Generated by Proof Engine v1.3.1.

提供机构:
Zenodo
创建时间:
2026-04-09
二维码
社区交流群
二维码
科研交流群
商业服务