RxSafeBench
收藏资源简介:
RxSafeBench是一个用于评估大型语言模型在模拟咨询场景中药物安全能力的综合基准数据集。该数据集由中国科学院深圳先进技术研究院的研究团队创建,包含2443个高质量的咨询场景,涵盖了禁忌症和药物相互作用类型。数据集通过模拟现实咨询对话,嵌入相关药物风险,并采用两阶段筛选策略确保临床真实性和专业质量。RxSafeBench旨在解决当前大型语言模型在药物安全方面的关键挑战,并提供了改进其可靠性的见解。
RxSafeBench is a comprehensive benchmark dataset for evaluating the medication safety capabilities of large language models (LLMs) in simulated consultation scenarios. This dataset was developed by a research team from the Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, and includes 2,443 high-quality consultation scenarios covering contraindications and drug interaction types. The dataset simulates real-world clinical consultation dialogues, embeds relevant medication risks, and adopts a two-stage screening strategy to ensure clinical authenticity and professional quality. RxSafeBench aims to address the key challenges faced by current large language models in medication safety, and provides insights for improving their reliability.




