遇见数据集

Source-faithfulness routing policies for AI verification of realist context-mechanism-outcome extractions: a multi-arm diagnostic stress test (data, code, and adjudication packets)

收藏
Zenodo2026-06-05 更新2026-06-05 收录
官方服务:

资源简介:

Data, code, prompts, and adjudication artifacts supporting the manuscript "Source-faithfulness routing policies for AI verification of realist context-mechanism-outcome (CMO) extractions: a multi-arm diagnostic stress test" (submitted to JMIR AI). Contains 766 per-record multi-vendor LLM verifier outputs (Anthropic, OpenAI, Google) across the eight panels reported in the manuscript, the analysis scripts, the resolved human-adjudication ground truth and aggregate inter-rater agreement, and the pre-specification record — sufficient to reproduce every headline number, including the pooled operational specificity of 83/87 = 95.4% (Wilson 95% CI 88.8–98.2%). Code is licensed MIT; data, CSVs, prompts, and documentation under CC BY 4.0. The 80 blinded adjudication packets and rendered blinded-prompt archive (copyrighted source passages), the individual adjudicator labels and de-blinding key (participant confidentiality), the vendor-dispatch wiring, and the calibration/probe panels not analysed in the paper are withheld, available from the corresponding author on reasonable request; none is required to reproduce a reported number. AMENDMENT 5 June 2026: A post-publication audit of this version identified two issues affecting the C02 arm of the A2 multi-arm validation, which in this deposit is labelled "Long 2022." First, the C02 source corpus was misidentified and contaminated: the folder labelled "Long 2022" in fact holds Taylor et al. 2024, "Care Under Pressure 2" (CUP2; BMJ Qual Saf 2024;33:523–538, doi:10.1136/bmjqs-2023-016468), and several unrelated source files (legacy hospital-implementation studies and other off-topic papers) had been mixed into the C02 record pool. Second, a programme-wide eligibility check that required each supporting quote to be exactly recoverable in its source after whitespace normalisation was found to discard the large majority of genuinely source-faithful quotes on OCR-derived PDFs, deflating the eligible-record pools across all arms. Both issues are being corrected. The corrective actions, all specified blind, before any new model output was generated are: (i) a verified-clean, SHA-256-frozen CUP2-only C02 corpus; (ii) a validated robust quote-grounding matcher (0% false-positive rate against a wrong-source negative control; ~85% recall versus ~11% for the previous exact-substring check); and (iii) a rebalanced five-arm specificity design (40 independent, human-adjudicated clean controls per arm). These corrections concern data integrity and accurate source attribution; they do not alter the locked decision policy, the validation prompt, the model set, or the four pre-specified estimands. Consequently, the C02 / "Long 2022" data and the corresponding notes in zenodo_metadata.json contained in THIS version are superseded and should not be used; they are retained only as the record under correction. Producing the corrected C02 results requires independent human re-adjudication of the controls and a re-run of the three-model validation, which is in progress. A corrected dataset will be deposited as a new version of this record. The full pre-specified amendment accompanies this deposit as PROTOCOL_AMENDMENT_2026-06-05.md.

提供机构:
Zenodo
创建时间:
2026-06-04
二维码
社区交流群
二维码
科研交流群
商业服务