遇见数据集

CVRT Structural Scoring of 358 Erdős Problem Statements: A Pre-Registered Blind Baseline with Within-Model Cross-Session Agreement, and Proposed Revisions to the CVRT-AI Hypothesis

收藏
Zenodo2026-06-12 更新2026-06-17 收录
官方服务:

资源简介:

This dataset freezes a blind structural scoring of 358 Erdős problem statement rows (240 unique problems) using the CVRT rubric v22.2 (five indicators: d_cat, d_sub, Gap II, R, B). Scoring was performed without reference to resolution status, with documented leakage exceptions. The deposit contains: (1) the frozen Rater-1 blind scores with per-problem narratives; (2) an independent fresh-session Rater-2 blind scoring of the same input; (3) agreement statistics between the two ratings — Gap II κw = 0.65, B κw = 0.53, d_sub κw = 0.62, R κw = 0.48, and d_cat κw = 0.36, the last falling below the pre-specified 0.40 threshold, triggering a pre-registered downgrade of all cross-rater d_cat claims; (4) post-hoc resolution flags (including the 9 problems solved by AlphaProof Nexus, May 2026) and excerpt-leakage flags with sensitivity analyses; (5) a retrospective 100-problem calibration reference (50 solved / 50 unsolved famous problems); and (6) proposed revisions to CVRT Theory Note v23: restating Heuristic 1 in two parts (proximity: low Gap II / low B; solver type: d_cat) and withdrawing its unsupported "high B" component, plus a unification of the conflicting Class definitions (Class = d_cat-based; point bands renamed S-band; totalling convention fixed as T-score v2). Both ratings were produced by the same model (Claude Fable 5, Anthropic) in separate sessions with no shared context; agreement statistics therefore measure within-model cross-session reproducibility, not generalizability across model families or human raters. All statistics are exploratory. This deposit is a timestamped pre-registered baseline and hypothesis-revision proposal, not a verification result: its primary purpose is to allow future AI and human resolution reports to be checked against these frozen scores. Problem statements are not redistributed; rows are identified by erdosproblems.com ID and URL. 本データセットは、CVRTルーブリックv22.2(5指標:d_cat, d_sub, Gap II, R, B)によるErdős問題358行(一意240問)の盲検構造採点を凍結したものである。採点は解決状況を参照せずに実施した(リーク例外は文書化済み)。収録物:(1) 第1採点の凍結スコアと根拠記述、(2) 別セッションによる独立第2盲検採点、(3) 両採点の一致統計——Gap II κw=0.65、B κw=0.53、d_sub κw=0.62、R κw=0.48、d_cat κw=0.36(事前指定閾値0.40未満のため、d_catに関する採点者間主張は単一採点者観察に格下げ)、(4) 事後付与の解決フラグ(AlphaProof Nexusの解決9問を含む)とリークフラグおよび感度分析、(5) 後ろ向き較正参照としての著名問題100問リスト(解決50・未解決50)、(6) CVRT理論ノートv23への修正提案——Heuristic 1の二部化(接近度:低Gap II・低B/解き手型:d_cat)と「高B」成分の撤回、およびClass定義の統一(Class=d_cat基準、点数区分はS-bandへ改称、合計規約T-score v2の固定)。 両採点は同一モデル(Claude Fable 5, Anthropic)の文脈を共有しない別セッションによるものであり、一致統計はモデル内セッション間再現性を測るものであって、異系統モデルや人間採点者への一般化可能性を示すものではない。統計はすべて探索的である。本資料は検証結果ではなく、タイムスタンプ付き事前登録基線および仮説修正提案であり、主目的は将来のAI・人間による解決報告を凍結スコアと照合可能にすることにある。問題文は再配布せず、各行はerdosproblems.comのIDとURLで識別される。

提供机构:
Zenodo
创建时间:
2026-06-12
二维码
社区交流群
二维码
科研交流群
商业服务