遇见数据集

TheoDB/french-pii-eval

收藏
Hugging Face2026-04-27 更新2026-05-03 收录
官方服务:

资源简介:

French PII Evaluation Dataset是一个专门用于法语个人身份信息(PII)检测评估和训练的数据集,旨在为TheoDB/privacy-filter-fr模型提供基准测试。数据集包含法语和英语内容,主要用于token-classification任务,涉及pii、privacy和french等标签。数据集规模在10K到100K之间,具体包括训练集(57,248个例子)、验证集(500个例子)和测试集(2,500个例子),以及一个英语回归检查集(426个例子)。数据集定义了8个PII类别,每个类别在测试集中至少有100个跨度。数据格式为OPF JSONL,兼容`opf eval`。数据来源于三个不同的数据集,经过去重和分层处理,确保测试集的严格独立性。

The French PII Evaluation Dataset is a curated dataset for evaluating and training French Personally Identifiable Information (PII) detection, specifically designed for benchmarking the TheoDB/privacy-filter-fr model. The dataset includes content in both French and English, targeting token-classification tasks with tags such as pii, privacy, and french. The dataset size ranges between 10K and 100K, comprising a training set (57,248 examples), a validation set (500 examples), and a test set (2,500 examples), along with an English regression check set (426 examples). It defines 8 PII classes, each with at least 100 spans in the test set. The data is formatted in OPF JSONL, compatible with `opf eval`. The dataset sources include three different datasets, which were deduplicated and stratified to ensure the test sets strict independence.

提供机构:
TheoDB
二维码
社区交流群
二维码
科研交流群
商业服务