Source-faithfulness routing policies for AI verification of realist context-mechanism-outcome extractions: a multi-arm diagnostic stress test (data, code, and adjudication packets)
收藏资源简介:
Data, code, prompts, and adjudication artifacts supporting the manuscript "Source-faithfulness routing policies for AI verification of realist context-mechanism-outcome (CMO) extractions: a multi-arm diagnostic stress test" (submitted to JMIR AI). Contains 766 per-record multi-vendor LLM verifier outputs (Anthropic, OpenAI, Google) across the eight panels reported in the manuscript, the analysis scripts, the resolved human-adjudication ground truth and aggregate inter-rater agreement, and the pre-specification record — sufficient to reproduce every headline number, including the pooled operational specificity of 83/87 = 95.4% (Wilson 95% CI 88.8–98.2%). Code is licensed MIT; data, CSVs, prompts, and documentation under CC BY 4.0. The 80 blinded adjudication packets and rendered blinded-prompt archive (copyrighted source passages), the individual adjudicator labels and de-blinding key (participant confidentiality), the vendor-dispatch wiring, and the calibration/probe panels not analysed in the paper are withheld, available from the corresponding author on reasonable request; none is required to reproduce a reported number.



