fuasfgauighsudghaughdoaughsdughdasughoadhg/pii-masking-300k
收藏资源简介:
Ai4Privacy PII 300k数据集是世界上最大的开放隐私掩码数据集,专门用于训练和评估模型以从文本中移除个人可识别信息(PII)和敏感信息,特别是在AI助手和大型语言模型(LLM)的背景下。该数据集包含两个子集:OpenPII-220k和FinPII-80k。OpenPII-220k涵盖27个PII类别(如用户名、时间等),针对749个讨论主题,覆盖教育、健康和心理学领域;FinPII-80k则额外包含约20个类别,专注于保险和金融。数据集总规模为30.4百万个文本标记,约220,000+个示例,其中包含7.6百万个PII标记。支持6种语言:英语(英国和美国)、法语(法国和瑞士)、德语(德国和瑞士)、意大利语(意大利和瑞士)、荷兰语(荷兰)和西班牙语(西班牙),并在8个司法管辖区进行了本地化。数据通过专有算法合成生成,无隐私侵犯风险,并经过人工循环验证,质量高(随机样本的标记标签准确率约98.3%)。训练/验证拆分为79%-21%。每个数据行包括原始文本、掩码后文本、隐私掩码标签、跨度标签、多语言BERT标记等信息,适用于令牌分类、文本生成等多种机器学习任务。
The Ai4Privacy PII 300k Dataset is the worlds largest open dataset for privacy masking, designed to train and evaluate models for removing personally identifiable information (PII) and sensitive data from text, particularly in the context of AI assistants and LLMs. It consists of two subsets: OpenPII-220k and FinPII-80k. OpenPII-220k includes 27 PII classes targeting 749 discussion subjects across education, health, and psychology; FinPII-80k adds approximately 20 additional classes tailored to insurance and finance. The dataset totals 30.4 million text tokens in approximately 220,000+ examples, with 7.6 million PII tokens. It supports 6 languages: English (UK and USA), French (France and Switzerland), German (Germany and Switzerland), Italian (Italy and Switzerland), Dutch (Netherlands), and Spanish (Spain), with strong localization across 8 jurisdictions. Data is synthetically generated using proprietary algorithms without privacy violations, and is human-in-the-loop validated for high quality (with ~98.3% token label accuracy on a random sample). The training/validation split is 79%-21%. Each row contains fields such as source text, target text, privacy mask labels, span labels, multilingual BERT tokens, etc., making it suitable for various machine learning tasks like token classification and text generation.



