遇见数据集

gravitee-io/pii-detection-dataset

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

Gravitee PII Detection是一个用于微调编码器风格PII(个人身份信息)和NER(命名实体识别)模型的统一多源语料库。它包含25个规范PII类别,提供字符级跨度标注,涵盖175,881条英文示例和781,052个实体跨度。数据集以单一训练分割发布,预期用于与外部PII语料库进行留出评估。数据来源包括多个上游数据集,如beki/privy、gretelai/gretel-pii-masking-en-v1等,覆盖对话、金融、结构化表单等领域的PII。标签方案包括AGE、PERSON、LOCATION、DATE_TIME等25个类别,支持英语文本中的PII检测,主要用于工作场所消息、结构化表单和金融对话。已知限制包括来源不平衡、合成数据主导、无多令牌/嵌套实体、仅限英语以及可能存在的领域转移。

Gravitee PII Detection is a harmonized, multi-source corpus for fine-tuning encoder-style PII (Personally Identifiable Information) and NER (Named Entity Recognition) models. It features 25 canonical PII classes, character-level span annotations, 175,881 English examples, and 781,052 entity spans. The dataset is published as a single split (train), with hold-out evaluation expected to be performed against unrelated external PII corpora. It integrates data from multiple upstream sources, such as beki/privy, gretelai/gretel-pii-masking-en-v1, and others, covering conversational, financial, and structured-form PII. The label scheme includes 25 classes like AGE, PERSON, LOCATION, and DATE_TIME, intended for PII detection in English text, particularly in workplace messages, structured forms, and financial conversations. Known limitations include source imbalance, dominance of synthetic data, absence of multi-token or nested entities, English-only support, and potential domain shift.

提供机构:
gravitee-io
二维码
社区交流群
二维码
科研交流群
商业服务