Data set for the Article: Rater choice determines measured fidelity: a multi-model evaluation of large language model raters for AI-generated plain-language biomedical summaries
收藏资源简介:
Data and analysis code for a study on large language model raters ("LLM-as-judge") applied to AI-generated plain-language biomedical summaries. Fifty peer-reviewed articles covering five conditions (alopecia areata, systemic lupus erythematosus, inflammatory bowel disease, diabetes, pulmonary disease) were summarized for four audiences. This produced 200 sections, of which 196 were analyzable. Four rater models, Claude Sonnet 4.6, DeepSeek V3.2, Llama 3.3-70B, and GPT-4o, scored each section using an 11-class source-anchored fidelity rubric and a fixed checklist of each article's findings, limitations, and safety statements. This resulted in 784 ratings and 3,565 claim-level judgments. One human rater validated a subsample of 10 articles (40 sections) based on condition. Running reproduce.py regenerates every number, table, and figure in the article without needing API credentials and at no cost. The source article PDFs are not redistributed; all 50 are referenced by DOI and PMID.



