SHIELD (Synthetic Human-annotated Identifier-replaced Entries for Learning and De-identification)
收藏资源简介:
SHIELD是由斯坦福大学医学技术数字解决方案团队构建的多样化临床笔记数据集,包含1,394条经过人工标注的临床文本,涵盖9类受保护健康信息(PHI)的10,505个标注片段。该数据集采用集合覆盖算法进行多样性采样,覆盖人口统计学和文档类型等多维度特征,并通过人机协同标注流程确保标注质量。数据集通过密码学替代技术处理原始PHI信息,既保护隐私又保留临床文本的语言结构特征,主要用于医疗记录去标识化研究,旨在解决传统基准数据集语义多样性不足、跨机构泛化性差等问题,为临床自然语言处理任务提供高质量评估基准。
SHIELD is a diverse clinical note dataset constructed by the Digital Solutions for Medical Technology Team at Stanford University. It contains 1,394 manually annotated clinical texts, including 10,505 annotated spans across 9 categories of Protected Health Information (PHI). This dataset employs set cover algorithm for diversity-driven sampling, covering multi-dimensional features such as demographics and document types, and guarantees annotation quality through a human-AI collaborative annotation pipeline. The dataset uses cryptographic substitution techniques to process the original PHI information, while safeguarding patient privacy and retaining the linguistic structural characteristics of clinical texts. It is primarily utilized for medical record de-identification research, and is designed to address the limitations of traditional benchmark datasets, such as inadequate semantic diversity and poor cross-institutional generalization, thereby providing a high-quality evaluation benchmark for clinical natural language processing tasks.
SHIELD 数据集概述
基本信息
- 数据集名称:SHIELD(A Diverse Clinical Note Dataset for Enterprise-Scale De-identification)
- 来源机构:斯坦福大学医学院(Stanford Medicine)
- 发布平台:Stanford Redivis(即将发布)
- 相关论文:arXiv:2605.03301
数据集规模与内容
- 临床笔记数量:1,394 份
- PHI(受保护健康信息)标注数量:10,505 个黄金标准跨度
- PHI 类别数量:9 个类别
- 构建方法:基于集合覆盖多样性抽样(set-cover diversity sampling),经过人工裁决
数据集特点
- 多样性:SHIELD 是多样化的临床笔记数据集,克服了 i2b2 2006/2014 等旧基准数据集过时且缺乏多样性的问题
- 独特性:使用 Fréchet Text Distance 和 Jensen-Shannon Divergence 进行的分布分析证实,SHIELD 在生物医学嵌入和词汇空间上占据与旧基准数据集不同的独特区域
数据获取方式
数据集将通过斯坦福大学医学院的 Redivis 平台公开发布,获取条件包括:
- 使用机构邮箱注册 Redivis 账户
- 签署数据使用协议
引用信息
bibtex @article{posada2025shield, title={SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification}, author={Posada, Jose D. and Love, David and Datta, Somalee and Desai, Priya}, journal={arXiv preprint arXiv:2605.03301}, year={2025} }
联系方式
- Jose D. Posada:jdposada@stanford.edu

- 1SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification斯坦福大学·医学技术数字解决方案 · 2026年



