遇见数据集

auren-research/pii-shield

收藏
Hugging Face2026-05-17 更新2026-05-31 收录
官方服务:

资源简介:

PII Shield 是一个大规模、多语言的数据集,专门用于训练和评估个人可识别信息(PII)检测模型。该数据集由Auren Research构建,结合了来自多个领域(包括企业邮件、法律、金融、医疗和通用领域)的真实世界文档,并提供了高质量的跨度级PII标注。标注是通过fastino/gliner2-privacy-filter-PII-multi模型完成的,该模型在SPY基准测试中取得了最高的F1分数,是开源PII检测器中的最佳表现。数据集旨在支持GDPR、LGPD、CCPA等相关隐私法规的生产级合规工作流。它包含约260万条示例,覆盖六种语言:英语、葡萄牙语、西班牙语、法语、德语和阿拉伯语,其中英语部分约53.1万条文档,其余为翻译示例。数据集涵盖40多种PII实体类型,如人名、联系方式、政府ID、银行信息、数字身份、秘密凭证和敏感日期等。数据来源于多个公开数据集,包括Enron Email Dataset、AI4Privacy、Gretel PII Masking和Nemotron-PII,并经过去重处理。英文文档通过标注模型进行标注后,使用占位符保留的翻译管道(基于tencent/HY-MT1.5-1.8B模型)翻译成其他语言,以确保PII实体在翻译过程中不被破坏。数据集结构包括按语言分区的Parquet文件,每个文件包含文本、PII检测状态、PII数量、PII类型和跨度信息。预期用例包括训练PII检测模型、基准测试多语言PII检测器、数据治理工具(如编辑和去标识化)、跨语言隐私信息提取研究以及微调编码器模型。数据集在CC BY 4.0许可证下发布,但使用时需注意标注和翻译的局限性,以及伦理考虑,例如不应将数据集用于监控或识别个人。

PII Shield is a large-scale, multilingual dataset for training and evaluating Personally Identifiable Information (PII) detection models. Built by Auren Research, it combines real-world documents from diverse domains (corporate email, legal, financial, healthcare, general) with high-quality span-level PII annotations produced by fastino/gliner2-privacy-filter-PII-multi — achieving the highest F1 on the SPY benchmark among open-source PII detectors. The dataset is designed to support production-grade compliance workflows under GDPR, LGPD, CCPA, and related privacy regulations. It contains approximately 2.6 million examples across six languages: English, Portuguese, Spanish, French, German, and Arabic, with about 531k English documents and the rest as translations. The dataset covers over 40 PII entity types, including person names, contact information, government IDs, banking details, digital identities, secrets, and sensitive dates. Source data is collected from public datasets such as Enron Email Dataset, AI4Privacy, Gretel PII Masking, and Nemotron-PII, and deduplicated. English documents are annotated using the annotation model, then translated into other languages via a placeholder-preserving translation pipeline based on tencent/HY-MT1.5-1.8B to prevent PII corruption. The dataset structure includes Parquet files partitioned by language, each containing text, PII detection status, PII count, PII types, and spans. Intended use cases include training PII detection models, benchmarking multilingual PII detectors, data governance tooling (e.g., redaction and de-identification), research on cross-lingual privacy information extraction, and fine-tuning encoder models. Released under CC BY 4.0 license, with limitations in annotation and translation quality, and ethical considerations against using it for surveillance or identification purposes.

提供机构:
auren-research
二维码
社区交流群
二维码
科研交流群
商业服务