遇见数据集

Data and R code for the revised Diagnostics 4509005 Arabic musculoskeletal LLM benchmark

收藏
Zenodo2026-09-24 更新2026-10-01 收录
官方服务:

资源简介:

This archive contains the 64-question Arabic musculoskeletal benchmark, 312 model-generated responses, confirmed ratings from three clinicians before and after audit-informed separate re-review, final majority scores, the historical submitted two-assessor consensus, audit flag and rating-change records, and the full R analysis with native outputs and execution evidence. The primary comparison uses 64 matched questions and the post-audit three-assessor majority. Pre-audit majority and separate assessor ratings are sensitivity analyses; submitted consensus is historical only. The archive preserves scoring-stage distinctions and does not pool them as independent observations. No human patient records are included. Reporting limitations and source roles are described in the README and manuscript.

提供机构:
Zenodo
创建时间:
2026-09-24
二维码
社区交流群
二维码
科研交流群
商业服务