遇见数据集

Data and code for: Which AI model is truly the best: how reliable are public benchmarks of frontier AI models?

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

Supplementary data and code for the paper Which AI model is truly the best: how reliable are public benchmarks of frontier AI models? (Kankaras, 2026; preprint DOI: 10.5281/zenodo.22998205). Can today's benchmarks tell the frontier models apart? On GPQA, Gemini 3 Pro's 1.1-point lead over GPT-5 is not a real difference, and a lead needs about 4 to 5 points before it is; on SWE-bench Verified, with the agent held fixed, Claude Opus 4.5 and Opus 4.6 cannot be told apart, and across the leaderboard none of the ten pairs among the top five submissions is a real difference. Among the newest models, on tau2-bench, Claude Opus 5, Grok 4.5 and GPT-5.6 lie within 1.7 points and none of the leads is real, while one model run on two builds of that benchmark's harness scores 25.3 and 40.2 per cent. flagship_summary.json holds these head-to-head verdicts for the flagship models on GPQA, MMLU-Pro, SWE-bench Verified, tau2-bench and MMLU. One psychometric report card per open AI benchmark whose model-by-item results are publicly downloadable. Each card reports, from the benchmark's own published response matrix: every model's score with an exact 95 per cent interval; the share of published scores whose interval includes or falls below the guessing floor; the standard error each score carries from item sampling; the number of distinguishable groups of models the ranking resolves, with multiplicity controlled across all pairwise comparisons; the concentration of discriminating variance across items; and, for suites that put the same items in more than one language, the share of items showing cross-language differential item functioning against a parametric bootstrap null. Any measure the published data cannot support is recorded as not computed with the reason, and every source examined and not carded is listed with the reason in skipped.json. This release holds 61 cards drawn from Desai et al. 2026, What AI Benchmarks Actually Measure; HELM (Stanford CRFM); Open LLM Leaderboard v1 archive; SWE-bench (Princeton and Stanford); and tau2-bench (Sierra). The method is the one set out in Ranks without resolution (doi:10.5281/zenodo.22128037), applied here to 61 public leaderboards. The record is a snapshot of 26 September 2026.

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务