BlueZeros/EHR-Ins-Reasoning
收藏资源简介:
EHR-Ins是一个为了增强大型语言模型在电子健康记录(EHR)上的推理和分析能力而开发的大规模、全面的指令数据集。它包含两种主要类型的数据:30万个高质量的推理案例和大约350万到400万个非推理案例。数据集覆盖了42种不同的EHR任务,分为决策制定(如诊断和治疗建议)和风险预测(如死亡率和再入院率)两大类。该数据集的核心创新是一个思维图驱动框架,用于大规模合成高质量推理数据。该框架通过识别EHR中的关键相关医疗实体,将这些实体与外部知识(如UMLS知识库)相链接,并促使模型(如GPT-4o)基于生成的图产生结构化的逐步临床推理。EHR-Ins提供了明确的医疗推理监督,使EHR-R1系列模型能够系统地获取进行准确和强大EHR分析所需的多样化、上下文丰富的推理能力。
EHR-Ins is a large-scale, comprehensive instruction dataset developed to enhance the reasoning and analysis capabilities of Large Language Models (LLMs) for Electronic Health Records (EHR). It consists of two major types of data: 300K high-quality reasoning cases and approximately 3.5 to 4 million non-reasoning cases. The dataset spans a wide variety of 42 distinct EHR tasks, categorized into two types: decision-making (e.g., diagnosis and treatment recommendations) and risk-prediction (e.g., mortality and readmission). The core innovation of the dataset is a thinking-graph-driven framework used to synthesize the high-quality reasoning data at scale. This framework works by identifying key related medical entities from EHRs, linking these entities with external knowledge such as the UMLS knowledge base, and prompting a model (like GPT-4o) to produce structured, step-by-step clinical reasoning based on the generated graph. EHR-Ins provides explicit medical reasoning supervision, enabling models like the EHR-R1 series to systematically acquire diverse, context-rich reasoning capabilities necessary for accurate and robust EHR analysis.




