遇见数据集

A determinacy audit of SWE-bench Pro: which tasks are underdetermined, with receipts

收藏
Zenodo2026-06-17 更新2026-06-21 收录
官方服务:

资源简介:

A preregistered determinacy audit of all 728 public SWE-bench Pro tasks. Each task is labeled for whether the behavior the hidden test grades is pinned by what a solver actually receives (the prose: problem + requirements + interface) plus the repository source. The decisive evidence is the repository itself: a codebase-determinacy sweep settles each prose-silent graded behavior against grep-verified live precedents at the base commit — one codebase way and gold matches it is determined-codebase (78 tasks, not a defect); one way and gold pins a different value is misdetermined; ≥2 live ways is codebase-plural; no comparable precedent and an absent constant is airtight. The claimable result is two-tier: a mechanical spine of 83 tasks (11.4%) provable from committed receipts by grep with no methodological buy-in (absent constants, contradicted conventions, multi-way choices, officially-graded prose-faithful alternative patches, a hand-verified clause), plus 26 two-expert reading-judgment splits (3.6%, cumulative 15.0%) where the prose itself licenses two faithful readings the hidden test splits. The prose-plurality splits are two-model adversarially verified: GPT-5.5 constructs the existence proof, an independent cross-family refuter (Claude opus) tries to kill it, a symmetric advocate pass recovers no missed splits (Cohen's κ = 0.52); the grep-mechanical tiers need no such pass. Three KNOWN_BAD gold-fails-grader defects and at least one KNOWN_MISMATCH (prose describes one feature, gold and test grade another) are recorded separately. Verdicts are mechanical and re-derivable from per-case receipts.

提供机构:
Zenodo
创建时间:
2026-06-17
二维码
社区交流群
二维码
科研交流群
商业服务