hivetrace/pii-bench
收藏资源简介:
PII-Bench (ru) 是一个用于评估俄语文本中个人数据(PII)检测质量的基准数据集。它采用span-level标注,明确标注实体的起始、结束字符索引和实体类型,旨在全面评估NER(命名实体识别)流程,包括ML模型和正则表达式,而不仅仅是IO、BIO或BILOU格式的模型评估。数据集包含13种实体类型(如姓名、电话号码、电子邮件地址、物理地址、银行卡号、CVC代码、INN、KPP、OGRN、OGRNIP、SNILS、护照号码和令牌),并分为9个域,包括敏感域(如银行、电信、交付、汽车、HR、房地产和一般支持)和低敏感域(如聊天和对话)。数据完全合成,使用Claude 4.5 Sonnet生成,不包含真实用户数据,仅用于研究、评估和基准测试目的,禁止用于模型训练或商业用途。数据集总计1810个示例,其中79%包含PII,字符级别PII占比约19.6%。
PII-Bench (ru) is a benchmark dataset for evaluating the quality of Personally Identifiable Information (PII) detection in Russian-language texts. It adopts span-level annotation, which explicitly marks the start and end character indices of entities as well as entity types, aiming to comprehensively evaluate Named Entity Recognition (NER) workflows including ML models and regular expressions, rather than only assessing models tagged with IO, BIO or BILOU formats. The dataset covers 13 entity types, such as full name, phone number, email address, physical address, bank card number, CVC code, INN, KPP, OGRN, OGRNIP, SNILS, passport number and token, and is divided into 9 domains, including sensitive domains (e.g., banking, telecommunications, delivery, automotive, HR, real estate and general support) and low-sensitivity domains (e.g., chat and conversation). All data in the dataset is synthetically generated using Claude 4.5 Sonnet, contains no real user data, is exclusively intended for research, evaluation and benchmarking purposes, and is prohibited from being used for model training or commercial purposes. In total, the dataset consists of 1810 instances, 79% of which contain PII, with the character-level proportion of PII reaching approximately 19.6%.



