遇见数据集

MedBias: A Counterfactual Demographic-Robustness Audit of LLM Clinical Decision Support — Data and Code

收藏
Zenodo2026-06-20 更新2026-06-21 收录
官方服务:

资源简介:

This repository contains the data, code, and stimulus set for the study "Auditing Demographic Robustness in Large Language Model Clinical Decision Support: A Counterfactual Probing Methodology." The project provides a reusable methodology for auditing whether large language models give different clinical recommendations when only a patient's demographic identity changes while all clinical facts are held constant. It includes 40 clinical vignette templates, a demographic-variant expander, a mechanical validity checker enforcing the holding-constant rule, a model-query harness, a structured-output parser, the full statistical-analysis and figure-generation scripts, the parsed model responses, and the complete results table (168 paired statistical tests with Holm-corrected p-values, bootstrap confidence intervals, and effect sizes). Across four current frontier models (GPT-4o-mini, GPT-4o, Claude Haiku, and Claude Sonnet), three demographic axes (sex, race/ethnicity, and socioeconomic status via insurance type), and 7,120 model decisions, no demographic comparison was statistically significant after correction, while substantial cross-model variation in clinical aggressiveness was observed.

提供机构:
Zenodo
创建时间:
2026-06-20
二维码
社区交流群
二维码
科研交流群
商业服务