A determinacy audit of SWE-bench Pro: which tasks are underdetermined, with receipts
收藏资源简介:
A preregistered determinacy audit of all 728 public SWE-bench Pro tasks. Each task is labeled for whether the behavior the hidden test grades is pinned by what a solver actually receives (the prose: problem + requirements + interface) plus the repository source. The decisive evidence is the repository itself: a codebase-determinacy sweep settles each prose-silent graded behavior against grep-verified live precedents at the base commit — one codebase way and gold matches it is determined-codebase (78 tasks, not a defect); one way and gold pins a different value is misdetermined; ≥2 live ways is codebase-plural; no comparable precedent and an absent constant is airtight. The claimable result is two-tier: a mechanical spine of 83 tasks (11.4%) provable from committed receipts by grep with no methodological buy-in (absent constants, contradicted conventions, multi-way choices, officially-graded prose-faithful alternative patches, a hand-verified clause), plus 26 two-expert reading-judgment splits (3.6%, cumulative 15.0%) where the prose itself licenses two faithful readings the hidden test splits. The prose-plurality splits are two-model adversarially verified: GPT-5.5 constructs the existence proof, an independent cross-family refuter (Claude opus) tries to kill it, a symmetric advocate pass recovers no missed splits (Cohen's κ = 0.52); the grep-mechanical tiers need no such pass. Three KNOWN_BAD gold-fails-grader defects and at least one KNOWN_MISMATCH (prose describes one feature, gold and test grade another) are recorded separately. Verdicts are mechanical and re-derivable from per-case receipts.



