bigbio/n2c2_2006_deid
收藏官方服务:
资源简介:
n2c2 2006 De-identification数据集来源于Partners Healthcare,仅包含医疗出院摘要。数据通过自动系统和手动验证两阶段进行标注,所有真实的个人健康信息(PHI)被替换为现实的替代品。数据集的任务是命名实体识别(NER)。数据集不公开,且未在PubMed中发布。
The n2c2 2006 De-identification Dataset is sourced from Partners Healthcare and only contains medical discharge summaries. The data was annotated through a two-stage process incorporating automated systems and manual verification, with all authentic protected health information (PHI) replaced by realistic substitutes. The task of this dataset is named entity recognition (NER). This dataset is not publicly available and has not been published in PubMed.
提供机构:
bigbio原始信息汇总
数据集概述
基本信息
- 名称: n2c2 2006 De-identification
- 语言: 英语
- 许可: DUA
- 多语言性: 单语种
详细描述
- 主页: https://portal.dbmi.hms.harvard.edu/projects/n2c2-nlp/
- 是否公开: 否
- 是否可在PubMed上获取: 否
- 任务: 命名实体识别(NER)
数据来源与处理
- 数据来源于Partners Healthcare,仅包含医疗出院总结。
- 数据通过标注和替换所有真实的个人健康信息(PHI)为现实的替代品进行准备。
- 数据标注过程分为两个阶段:首先使用自动系统进行标记,随后由三名标注者手动验证并讨论标记结果,最终确定PHI标签。
数据结构
- 原始数据集不包含每个实体的跨度,跨度在加载器中计算,最终与标签对应的文本以源格式保存。



