遇见数据集

Paired English-Arabic benchmark of large language models on clinical records: data and code

收藏
Zenodo2026-09-25 更新2026-10-01 收录
官方服务:

资源简介:

Data and code accompanying the article "Where the Arabic penalty actually falls: a paired English-Arabic benchmark of large language models on clinical records". The benchmark presents the same synthetic patient record to six large language models in three registers: English, code-switched Arabic that keeps Latin-script medication and laboratory terms, and full Modern Standard Arabic. Two record-grounded tasks, medication reconciliation and discharge planning, are posed at two record-quality levels (clean and defect-injected). This gives 576 items in 192 item sets, with language as the only factor that varies within a set. Scoring is deterministic and identical across languages. The archive contains: the eight synthetic patient records generated with Synthea, with rule-derived gold standards; the item design; a 211-term English-Arabic clinical lexicon with its term-by-term back-translation audit; the Arabic record templates and task instructions; all 576 raw model responses with token usage and stop reasons; item-level and error-level scored datasets; the rendering, inference, scoring and harm-grading code (Python); the statistical analysis script (R); and a README with a data dictionary and reproduction instructions. Rescoring the raw responses with the included code, and running the analysis script, reproduce the published results exactly. SHA-256 checksums are provided for every file. All patient records are synthetic. The archive contains no real patient data, no human participants and no clinician ratings. Licences: data under CC BY 4.0; code under the MIT licence.

提供机构:
Zenodo
创建时间:
2026-09-25
二维码
社区交流群
二维码
科研交流群
商业服务