遇见数据集

Performance of Large Language Models as a Tool for Primary Care Consultations

收藏
Zenodo2026-03-20 更新2026-05-26 收录
官方服务:

资源简介:

Abstract We present two datasets. On the one hand, a set of health-related queries from real users. This dataset was extracted from the 2025 Google report and validated by a panel of physicians from different specialties. It represents the primary care queries most frequently made worldwide in the English language. On the other hand, we present a set of responses generated by different Large Language Models for these queries, annotated by a panel of physicians distinct from the previous one. These responses were evaluated across several dimensions, such as medical consensus, correctness, and potential harm. The models used to generate these responses were: an open model (Llama3, version llama3:8b-instruct-q4_0), a widely used proprietary model (GPT-4, version o-mini), and an open model specialized through clinical fine-tuning (MedLlama3, version llama3-med42-8b:latest). Citation information If you use this version of the dataset, please cite as follows: Fernández-Pichel, M., Pascual Presa, N., Losada, D. E., García Orosa, B., Gude, F., Costa Lathan, C., Sueiro Justel, J., Gómez Fontenla, A., Lastra Pérez, M., & Alonso García, F. (2026). Performance of Large Language Models as a Tool for Primary Care Consultations [Data set]. Zenodo. https://doi.org/10.5281/zenodo.18311708 Acknowledgements This dataset is part of the R&D project Artificial Intelligence in Digital Media in Spain: Effects and Roles ( PID2024-156034OB-C22), funded by MICIU/AEI/10.13039/ 501100011033 and by “ERDF/EU”. This research is also supported by the the project Cátedra de IA aplicada a la Medicina Personalizada de Precisión (Cátedras ENIA, TSI-100932-2023-3); Cátedras ENIA is funded by the Ministerio de Transformación Digital y Función Pública (Secretaría de Estado de Digitalización e Inteligencia Artificial); and by the NextGeneration EU-fund. The second and third author also thank the financial support from the Agencia Estatal de Investigación (Spain) (PID2022-137061OB-C22 funded by MICIU/AEI/10.13039/ 501100011033), the Xunta de Galicia - Conselleria de Educación, Ciencia, Universidades e Formación Profesional (Centro de investigación de Galicia acreditación 2024-2027 ED431G-2023/04 and Reference Competitive Group accreditation ED431C 2022/19) and the European Union (European Regional Development Fund - ERDF).

摘要 本研究发布两个数据集。其一为源自真实用户的健康相关查询数据集,该数据集提取自2025年谷歌报告,并经多专科医师团队验证,代表了全球范围内使用英语进行的最常见初级保健咨询问题。 其二为针对上述查询由多款大语言模型(Large Language Model, LLM)生成的回复数据集,由另一组独立于前述团队的医师团队进行标注。这些回复从医学共识、正确性及潜在危害等多个维度开展评估。用于生成回复的模型包括:一款开源模型(Llama3,版本号llama3:8b-instruct-q4_0)、一款广泛使用的专有模型(GPT-4,版本号o-mini),以及一款经过临床微调的专业开源模型(MedLlama3,版本号llama3-med42-8b:latest)。 引用信息 若使用本版本数据集,请按如下格式引用: Fernández-Pichel, M., Pascual Presa, N., Losada, D. E., García Orosa, B., Gude, F., Costa Lathan, C., Sueiro Justel, J., Gómez Fontenla, A., Lastra Pérez, M., & Alonso García, F. (2026). 作为初级保健咨询工具的大语言模型性能 [数据集]. Zenodo. https://doi.org/10.5281/zenodo.18311708 致谢 本数据集隶属于西班牙「数字媒体中的人工智能:影响与角色」研发项目(项目编号PID2024-156034OB-C22),该项目由西班牙科学创新与大学部/西班牙国家研究局(MICIU/AEI/10.13039/501100011033)以及欧洲区域发展基金(ERDF/EU)资助。本研究同时得到「面向精准个性化医学的人工智能应用教席」(Cátedra de IA aplicada a la Medicina Personalizada de Precisión,项目编号Cátedras ENIA, TSI-100932-2023-3)支持;Cátedras ENIA由西班牙数字化与智能功能转型部(数字化与人工智能国务秘书处)资助,并通过下一代欧盟基金提供支持。第二及第三作者感谢西班牙国家研究局(MICIU/AEI/10.13039/501100011033资助的项目PID2022-137061OB-C22)、加利西亚自治区教育、科学、大学与职业培训委员会(加利西亚研究中心认证2024-2027 ED431G-2023/04及竞争性参考团队认证ED431C 2022/19)以及欧盟(欧洲区域发展基金ERDF)提供的经费支持。

提供机构:
Zenodo
创建时间:
2026-03-20
二维码
社区交流群
二维码
科研交流群
商业服务