遇见数据集

Truth Contradiction-and-Correction Benchmark v1.0

收藏
Zenodo2026-07-20 更新2026-08-01 收录
官方服务:

资源简介:

A frozen, versioned, seed-reproducible multi-session benchmark for user-facing epistemic behavior in conversational agents: 120 seeded three-session conversations (24 turns each) per instance set, mixing asks, cross-session re-asks, user statements, and user corrections, with contradictions and false corrections injected by the generator; labels are the generator's injection record, never any system's output. Ships two frozen sets — v1.0 (synthetic 60-subject world, generation seed 20260712) and stageb-v1.0 (real 60-fact geography world, generation seed 20260713; 422 injected corrections per run: 229 true, 193 false) — each with instances.jsonl, labels.jsonl, and a MANIFEST.json carrying SHA-256 hashes and generation parameters, plus the seeded generator (build_benchmark.py) and a language-agnostic scorer (score.py) over JSONL event logs whose metric definitions are the spec: expression fidelity, unacknowledged self-contradiction, acknowledgment auditability, true/false correction acceptance, confidently-asserted-false rate, and coverage/accuracy. Constructed data only; it measures machine-scored expression-and-revision behavior on knowable cases, not believability, trust, open-domain consistency, or human perception. Accompanies the paper 'Truth for Believable AI: Expressed Doubt, Provenance, and Belief Revision as an Engineerable Stance'.

提供机构:
Zenodo
创建时间:
2026-07-20
二维码
社区交流群
二维码
科研交流群
商业服务