“启真”医学知识数据
收藏资源简介:
本数据对医疗场景中的决策具有重要指导意义,使得决策更加科学、更加符合医疗行业的规范,可作为医疗场景中人机交互的逻辑依据。本数据经过专家医生团队校验认证,具有权威性,数据范围包括疾病、药品、检查检验、手术、诊疗决策路径、健康宣教等。本数据可用于医疗人工智能产品的模型训练,例如临床辅助决策系统、医疗相关的大语言模型等产品的模型训练,经过本数据训练的模型能提供更加专业的建议结果。从医学权威机构官方渠道获取原始数据后,使用自然语言处理(NLP)从大量的医学文本数据中自动识别、抽取和整合有用的信息。首先,对原始文本进行预处理,包括分词(将句子分割成单词)、去除停用词(如“的”、“是”等无实际含义的词语)以及词干化(将单词还原为其词干形式),以减少数据噪音,使得文本更易于处理。然后,通过命名实体识别(Named Entity Recognition,NER)技术识别文本中的特定实体,如疾病、症状、药物、治疗方法等。其次,通过关系抽取技术从文本中提取实体之间的关系,如“疾病A可以通过药物B治疗”。通过以上算法规则初步得到了结构化的医学知识数据,然后通过医生专家团队的审核、校验及认证,形成高质量可用的医学知识数据。
This dataset provides critical guidance for decision-making in medical scenarios, enhancing the scientific rigor and compliance with medical industry standards of such decisions, and can serve as a logical foundation for human-computer interaction in medical scenarios. This dataset has been verified and certified by a team of expert physicians, endowing it with authoritative validity; its scope covers diseases, medications, examinations and laboratory tests, surgeries, diagnosis and treatment decision pathways, health education, and more. This dataset can be used for model training in medical artificial intelligence products, such as clinical decision support systems and medical-related large language models (LLMs); models trained with this dataset can deliver more professional advisory outcomes. After acquiring raw data from official channels of authoritative medical institutions, natural language processing (NLP) techniques are utilized to automatically identify, extract and integrate useful information from large volumes of medical textual data. First, preprocessing is conducted on the raw text, including tokenization (splitting sentences into individual words), stopword removal (removing meaningless Chinese function words such as "的" (equivalent to "the/of"), "是" (equivalent to "is/am/are"), and other semantically empty terms), and stemming (reducing words to their root forms) to reduce data noise and facilitate text processing. Subsequently, named entity recognition (NER) technology is employed to identify specific entities in the text, such as diseases, symptoms, medications, treatment methods, and so on. Next, relationship extraction technology is applied to extract the relationships between entities from the text, for example, "Disease A can be treated with Medication B". Structured medical knowledge data is initially obtained via the aforementioned algorithmic rules, and then reviewed, verified and certified by the expert physician team to form high-quality, usable medical knowledge datasets.




