遇见数据集

VerifiMind-PEAS Evaluation Dataset v1.0: Human-Annotated Ground Truth for Multi-Agent Epistemic Verification

收藏
Zenodo2026-07-10 更新2026-08-01 收录
官方服务:

资源简介:

A 100-item human-annotated evaluation dataset for measuring the performance of the VerifiMind-PEAS multi-agent epistemic verification system. Contains ground-truth labels assigned by a single domain-expert annotator (ANN-AL01) across 5 domains (science, technical, ethics, business, general). Includes system baseline predictions, annotator confidence scores, and detailed reasoning notes. Inter-rater agreement: Cohen's κ = 0.504 (moderate, below pre-registered 0.60 bar — reported honestly). Zero polar confusions (PROCEED↔REJECT). AI comprehension assistance disclosed voluntarily; influence audit finds no contamination evidence. Part of the Genesis Prompt Engineering Methodology validation framework.

提供机构:
Zenodo
创建时间:
2026-07-09
二维码
社区交流群
二维码
科研交流群
商业服务