VerifiMind-PEAS Evaluation Dataset v1.0: Human-Annotated Ground Truth for Multi-Agent Epistemic Verification
收藏官方服务:
资源简介:
A 100-item human-annotated evaluation dataset for measuring the performance of the VerifiMind-PEAS multi-agent epistemic verification system. Contains ground-truth labels assigned by a single domain-expert annotator (ANN-AL01) across 5 domains (science, technical, ethics, business, general). Includes system baseline predictions, annotator confidence scores, and detailed reasoning notes. Inter-rater agreement: Cohen's κ = 0.504 (moderate, below pre-registered 0.60 bar — reported honestly). Zero polar confusions (PROCEED↔REJECT). AI comprehension assistance disclosed voluntarily; influence audit finds no contamination evidence. Part of the Genesis Prompt Engineering Methodology validation framework.
提供机构:
Zenodo创建时间:
2026-07-09



