遇见数据集

PeterPaker123/mimic-iv-clinical-ner-aug

收藏
Hugging Face2026-05-07 更新2026-05-31 收录
官方服务:

资源简介:

这是一个医疗领域的命名实体识别数据集,专注于从医疗文本中识别关键实体。数据包含文本标记(tokens)和对应的命名实体标签(ner_tags),标签类别包括主诉、出院状况、出院诊断、出院指导、出院药物、现病史、影像学检查、入院药物、既往病史、相关结果、体格检查等医疗实体,并采用BIO标注格式(B-表示实体开始,I-表示实体内部,O-表示非实体)。数据集分为训练集(86243个示例)、验证集(21561个示例)和测试集(26952个示例),适用于医疗信息提取和自然语言处理任务。

This is a medical domain named entity recognition dataset focused on identifying key entities from medical texts. The data includes text tokens and corresponding named entity labels (ner_tags), with label categories such as chief complaint, discharge condition, discharge diagnosis, discharge instructions, discharge medications, history of present illness, imaging, medications on admission, past medical history, pertinent results, physical exam, etc., using the BIO annotation format (B- for entity beginning, I- for entity inside, O- for non-entity). The dataset is split into training set (86,243 examples), validation set (21,561 examples), and test set (26,952 examples), suitable for medical information extraction and natural language processing tasks.

提供机构:
PeterPaker123
二维码
社区交流群
二维码
科研交流群
商业服务