遇见数据集

Claim Verification: "The pattern-matching limitations identified in GSM-NoOp are practically surmountable when LLMs are allowed to offload formal reasoning steps to code execution." — Proved

收藏
Zenodo2026-04-08 更新2026-05-29 收录
官方服务:

资源简介:

Automated fact-verification of the claim: "The pattern-matching limitations identified in GSM-NoOp are practically surmountable when LLMs are allowed to offload formal reasoning steps to code execution." Verdict: PROVED Key Findings SC1 confirmed (4 of 3 required sources): The GSM-NoOp benchmark demonstrates that LLMs suffer catastrophic accuracy drops (up to 65 percentage points) when irrelevant but syntactically plausible information is added to math problems, confirming reliance on pattern matching rather than formal reasoning. SC2 confirmed (4 of 3 required sources): Program-aided approaches (PAL, PoT, IIPC) that offload computation to code execution achieve significant accuracy gains over chain-of-thought methods, doing so by providing "a deterministic path to solutions while minimizing calculation errors" — structurally bypassing the pattern-matching failure mode. All 8 citations verified across the source pages. Important caveat: No study has directly evaluated code-execution methods on the GSM-NoOp dataset. SC2 relies on mechanistic evidence — code execution forces explicit variable binding that structurally prevents the irrelevant-information integration failure GSM-NoOp exploits. Files proof.py — Re-runnable Python verification script proof.md — Structured proof report proof_audit.md — Full verification audit trail proof_narrative.md — Plain-language summary proof.json — Machine-readable structured data Generated by Proof Engine v1.10.0.

提供机构:
Zenodo
创建时间:
2026-04-08
二维码
社区交流群
二维码
科研交流群
商业服务