MA-CERD: A Gold Benchmark Dataset for Counterfactual Requirements and Evidence Change Reasoning in Multi-Agent Systems
收藏资源简介:
MA-CERD (Multi-Agent Constraint-and-Evidence Reasoning Dataset) is a Gold benchmark dataset for evaluating reasoning under counterfactual requirement and evidence changes in multi-agent systems. The release contains 1,000 adjudicated Gold records constructed from 100 base cases and atomic semantic interventions. Each record captures the required relation, handling decision, impact set, preservation set, and provenance required to assess whether a system responds correctly to a changed requirement or evidence condition while preserving unaffected elements. Gold construction combines three separated evidence layers: O1 explicit-rule evidence, O2 blind human semantic review with third-party adjudication of disagreements, and O3 deterministic simulator-based shadow execution. The final release includes schema definitions, data dictionaries, family-level splits, privacy checks, evaluation utilities, intended-use documentation, and SHA-256 integrity manifests. The benchmark is intended for research on requirements evolution, change-impact reasoning, evidence-aware decision support, multi-agent systems, and leakage-resistant evaluation. It does not constitute operational validation, certified simulator fidelity, deployment readiness, or real-world safety assurance.



