遇见数据集

Performance of Large Language Models as a Tool for Primary Care Consultations

收藏
Zenodo2026-03-20 更新2026-05-26 收录
官方服务:

资源简介:

Abstract We present two datasets. On the one hand, a set of health-related queries from real users. This dataset was extracted from the 2025 Google report and validated by a panel of physicians from different specialties. It represents the primary care queries most frequently made worldwide in the English language. On the other hand, we present a set of responses generated by different Large Language Models for these queries, annotated by a panel of physicians distinct from the previous one. These responses were evaluated across several dimensions, such as medical consensus, correctness, and potential harm. The models used to generate these responses were: an open model (Llama3, version llama3:8b-instruct-q4_0), a widely used proprietary model (GPT-4, version o-mini), and an open model specialized through clinical fine-tuning (MedLlama3, version llama3-med42-8b:latest). Citation information If you use this version of the dataset, please cite as follows: Fernández-Pichel, M., Pascual Presa, N., Losada, D. E., García Orosa, B., Gude, F., Costa Lathan, C., Sueiro Justel, J., Gómez Fontenla, A., Lastra Pérez, M., & Alonso García, F. (2026). Performance of Large Language Models as a Tool for Primary Care Consultations [Data set]. Zenodo. https://doi.org/10.5281/zenodo.18311708 Acknowledgements This dataset is part of the R&D project Artificial Intelligence in Digital Media in Spain: Effects and Roles ( PID2024-156034OB-C22), funded by MICIU/AEI/10.13039/ 501100011033 and by “ERDF/EU”. This research is also supported by the the project Cátedra de IA aplicada a la Medicina Personalizada de Precisión (Cátedras ENIA, TSI-100932-2023-3); Cátedras ENIA is funded by the Ministerio de Transformación Digital y Función Pública (Secretaría de Estado de Digitalización e Inteligencia Artificial); and by the NextGeneration EU-fund. The second and third author also thank the financial support from the Agencia Estatal de Investigación (Spain) (PID2022-137061OB-C22 funded by MICIU/AEI/10.13039/ 501100011033), the Xunta de Galicia - Conselleria de Educación, Ciencia, Universidades e Formación Profesional (Centro de investigación de Galicia acreditación 2024-2027 ED431G-2023/04 and Reference Competitive Group accreditation ED431C 2022/19) and the European Union (European Regional Development Fund - ERDF).

摘要 本研究构建了两类数据集。其一为源自真实用户的健康相关查询数据集,该数据集提取自2025年谷歌报告,经多专科医师团队审核验证,涵盖全球范围内英语语境下最常见的基层医疗咨询问题。 其二为针对上述查询,由不同大语言模型(Large Language Model)生成的回复数据集,该数据集经与前述团队不同的医师团队标注审核。这些回复从多个维度进行了评估,包括医学共识性、准确性以及潜在危害性。用于生成回复的模型包括:开源模型(Llama3,版本llama3:8b-instruct-q4_0)、广泛使用的闭源模型(GPT-4,版本o-mini),以及经临床微调的专业开源模型(MedLlama3,版本llama3-med42-8b:latest)。 引用信息 若使用本版本数据集,请按如下格式引用: Fernández-Pichel, M., Pascual Presa, N., Losada, D. E., García Orosa, B., Gude, F., Costa Lathan, C., Sueiro Justel, J., Gómez Fontenla, A., Lastra Pérez, M., & Alonso García, F. (2026). Performance of Large Language Models as a Tool for Primary Care Consultations [Data set]. Zenodo. https://doi.org/10.5281/zenodo.18311708 致谢 本数据集隶属于西班牙“数字媒体中的人工智能:影响与角色”研发项目(项目编号PID2024-156034OB-C22),该项目由MICIU/AEI/10.13039/501100011033以及欧盟区域发展基金(ERDF/EU)资助。本研究同时获得“面向精准个性化医疗的人工智能教席”(Cátedras ENIA,项目编号TSI-100932-2023-3)项目支持;该教席项目由西班牙数字化与公共职能转型部(数字化与人工智能国务秘书处)资助,并获得欧盟下一代复苏基金支持。第二及第三作者感谢西班牙国家研究署(Agencia Estatal de Investigación)的经费支持(项目PID2022-137061OB-C22,由MICIU/AEI/10.13039/501100011033资助)、加利西亚自治区教育、科学、大学与职业培训委员会(加利西亚研究中心认证2024-2027 ED431G-2023/04以及竞争性参考团队认证ED431C 2022/19),以及欧盟(欧洲区域发展基金ERDF)的支持。

提供机构:
Zenodo
创建时间:
2026-01-20
二维码
社区交流群
二维码
科研交流群
商业服务