lucianfialho/privacy-filter-br-dataset
收藏资源简介:
privacy-filter-br数据集是一个用于微调privacy-filter-br命名实体识别(NER)模型的合成葡萄牙语(巴西)个人身份信息(PII)检测数据集。该数据集包含22个PII类别,兼容BIOES标注格式。每个数据样本是一个JSON对象,包含完整文本(text)、实体标注列表(entities,包括起始位置、结束位置和标签)以及生成模板(template)。PII类别涵盖巴西结构化标识符(如CPF、CNPJ、RG等)、个人联系信息(如姓名、电子邮件、电话等)、B2B/电子商务相关标识(如订单ID、客户ID等)以及OAI分类(如URL、账号、密钥)。数据集通过合成配置文件、Jinja2模板、大型语言模型(LLM)重写和格式感知标注器生成,当前稳定版本为v8.1,包含172,075个训练文档和17,072个保留文档。数据仅包含合成信息,所有PII均为虚构但格式符合巴西规范,适用于隐私过滤任务,但需注意合成数据与真实数据之间的性能差距。
The privacy-filter-br Dataset is a synthetic Portuguese (Brazilian) PII detection dataset used to fine-tune the privacy-filter-br NER model. It includes 22 PII categories and is compatible with BIOES tagging. Each data sample is a JSON object containing the full text (text), a list of entity annotations (entities with start/end positions and labels), and the generation template (template). The PII categories cover Brazilian structured identifiers (e.g., CPF, CNPJ, RG), personal contact information (e.g., person name, email, phone), B2B/e-commerce identifiers (e.g., order ID, customer ID), and OAI taxonomy (e.g., URL, account number, secret). The dataset is generated through synthetic profiles, Jinja2 templates, LLM rewriting, and format-aware labeling. The current stable version is v8.1, with 172,075 training documents and 17,072 holdout documents. It is synthetic-only, with all PII being fictional but formatted according to Brazilian standards, suitable for privacy filtering tasks, though a synthetic-to-real performance gap should be considered.




