遇见数据集

anonupload1ng/toulmin_errors

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

Toulmin-Errors 是一个用于类型化推理错误检测的基准数据集。该数据集旨在解决当前推理错误基准大多只衡量事实和逻辑错误,而很少评估两种论证层面的失败:声称范围错误(Qualifier, Q)和忽略反证据(Rebuttal, R)。基于Toulmin的论证模型,错误被分类为四个维度:Grounds(前提/事实)、Warrant(推理步骤)、Qualifier(范围/条件)和Rebuttal(反证据)。数据集包含三个部分:1) 受控腐败集(medreason_qr_corruption),包含927个医疗推理链,这些链经过专家验证为正确,然后在已知步骤中注入一个已知的Q或R错误,每个案例有一个已知类型和位置的错误,并通过独立盲分类器验证;2) 自然错误池,包括来自AI代理的真实推理轨迹,这些轨迹得出了错误结论,没有注入错误,其中natural_errors_*包含82个材料科学可行性轨迹,evidence_inference_*包含30个临床随机对照试验推理轨迹;3) 对六个现有推理错误基准(如bigbench_typed、processbench_typed等)的类型化重新注释,将每个现有错误标签重新标记为Toulmin维度,总计约3,839个类型化错误。该数据集可用于研究Q和R错误,并展示这些错误在形式推理数据集中罕见,但在科学推理数据集中常见。

Toulmin-Errors: A Benchmark for Typed Reasoning-Error Detection. This benchmark provides data to study reasoning errors, focusing on two argument-level failures: getting the scope of a claim wrong (Qualifier, Q) and ignoring counter-evidence (Rebuttal, R). Based on Toulmins argument model, errors are typed along four dimensions: Grounds (premises/facts), Warrant (inferential step), Qualifier (scope/conditions), and Rebuttal (counter-evidence). The benchmark consists of three parts: 1) A controlled corruption set (medreason_qr_corruption, n=927) with expert-verified medical reasoning chains where one step is rewritten to inject a known Q or R error at a known location, verified by an independent blind classifier; 2) Natural-error pools, including real reasoning traces from AI agents that reached wrong conclusions without injected errors, with natural_errors_* containing 82 materials-science feasibility traces and evidence_inference_* containing 30 clinical-RCT reasoning traces; 3) Typed re-annotations of six existing reasoning-error benchmarks (bigbench_typed, processbench_typed, prm800k_typed, mrben_typed, deltabench_typed, and legalbench_typed as a negative control), where pre-existing labeled errors are re-tagged with Toulmin dimensions, totaling ~3,839 typed errors. The dataset shows that Q+R failures are rare in formal-reasoning datasets but common in scientific-reasoning ones.

提供机构:
anonupload1ng
二维码
社区交流群
二维码
科研交流群
商业服务