遇见数据集

lianghsun/tw-PII-bench

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

该数据集是一个用于评估个人身份信息(PII)检测器的标记分类基准,专注于台湾特有的繁体中文个人信息。它包含三种按文本长度划分的数据集(短、中、长),以测试模型在不同上下文长度下的表现。数据集涵盖了8种模型已有的标签类别和11种台湾特有的超出模型标签类别的PII,以及5种硬性负面样本。该基准旨在揭示PII检测器在台湾场景下的标签覆盖差距和特定地区的失败模式。所有PII均为虚构,避免使用真实个人信息。

This dataset is a token classification benchmark for evaluating Personally Identifiable Information (PII) detectors, focusing on Taiwan-specific Traditional Chinese PII. It includes three subsets categorized by text length (short, medium, long) to assess model performance across varying context lengths. The dataset covers 8 existing label categories from standard models, 11 Taiwan-specific PII categories beyond the standard model's label set, plus 5 hard negative samples. This benchmark aims to uncover the label coverage gaps and region-specific failure modes of PII detectors in Taiwanese application scenarios. All PII contained herein are entirely fictional, with no real personal information utilized.

提供机构:
lianghsun
二维码
社区交流群
二维码
科研交流群
商业服务