遇见数据集

MedConform-Bench: experimental data and code for medical LLM conformity evaluation

收藏
Zenodo2026-08-15 更新2026-08-20 收录
官方服务:

资源简介:

De-identified experimental records and analysis code for MedConform-Bench (PLOS ONE submission). This version (v3) corrects the record counts and adds the configuration files, environment specification, and audited analysis outputs used in the manuscript. Includes:- Main experiment: 40,500 model response records (6 models × 3 datasets × 5 protocols × 150 items × 3 technical repeats)- Mitigation experiment: 24,300 Wrong-Guidance records (baseline / CP-1 / CP-2 × 3 repeats)- Combined total: 64,800 responses- configs/models.yaml and configs/experiment.yaml (public identifiers and environment-variable names only; no credentials)- Audited exact-conformity tables and the standalone audit script (supporting/S3)- High-risk labeling records from two deterministic rule sets (153 items) and HRC sensitivity outputs (supporting/S2)- Manuscript figures Fig1–Fig9 redrawn from the audited definitions The six models are DeepSeek-V3, Qwen2.5-72B-Instruct, Qwen3-32B, HuatuoGPT-o1-7B, Apollo2-7B, and Med42-8B. Conformity in the manuscript is exact selection of the unanimous wrong agent-endorsed option among Raw-correct observations. High-risk labels were produced by automated rules; no human annotation was collected. Public benchmark items are from PubMedQA, MedQA-USMLE, and MedMCQA under their original licenses.

提供机构:
Zenodo
创建时间:
2026-08-15
二维码
社区交流群
二维码
科研交流群
商业服务