joelbarmettler/gheim-ch-pii-212k
收藏资源简介:
gheim-ch-pii-212k 是一个多语言个人身份信息(PII)命名实体识别数据集,包含212,503个文本块。数据集覆盖瑞士四种官方语言(德语、法语、意大利语、罗曼什语)和英语,专注于8种PII类别:账户号码、私人地址、私人日期、私人电子邮件、私人人物、私人电话、私人URL和秘密信息。数据集中84%为真实文本,来自Apertus预训练语料库(包括瑞士法院裁决、联邦议会记录、瑞士过滤的网络文本和罗曼什语语料库),16%为模板和LLM生成的合成文本,用于填补真实文本覆盖不足的类别。注释由三个独立的大型语言模型(Gemma 4 26B-A4B、Qwen3.6 35B-A3B、Nemotron-3 Nano Omni 30B-A3B)和带有校验和验证器的正则表达式目录生成,采用33类BIOES标记方案。数据集分为训练、验证和测试集,真实文本部分遵循文档级隔离,确保同一文档的文本块不会跨越不同分割。数据集基于CC BY 4.0许可证发布,适用于隐私保护和数据匿名化任务。
gheim-ch-pii-212k is a multilingual named entity recognition (NER) dataset focused on personally identifiable information (PII), consisting of 212,503 text chunks. The dataset covers four official languages of Switzerland (German, French, Italian, Romansh) plus English, and targets 8 PII categories: account numbers, private addresses, private dates, private emails, private persons, private phone numbers, private URLs, and sensitive information. 84% of the dataset comprises authentic text sourced from the Apertus pre-training corpus, which includes Swiss court rulings, federal parliamentary records, filtered Swiss web text, and the Romansh language corpus; the remaining 16% is synthetic text generated via templates and large language models (LLMs) to compensate for insufficient authentic text coverage across categories. Annotations were produced by three independent large language models (Gemma 4 26B-A4B, Qwen3.6 35B-A3B, Nemotron-3 Nano Omni 30B-A3B) along with a regex catalog equipped with checksum validators, adopting a 33-class BIOES tagging schema. The dataset is divided into training, validation, and test sets, with the authentic text subset following document-level isolation to prevent text chunks from the same document from being split across different subsets. Released under the CC BY 4.0 license, the dataset is applicable to privacy protection and data anonymization tasks.



