TheoDB/french-pii-eval
收藏资源简介:
French PII Evaluation Dataset是一个专门用于法语个人身份信息(PII)检测评估和训练的数据集,旨在为TheoDB/privacy-filter-fr模型提供基准测试。数据集包含法语和英语内容,主要用于token-classification任务,涉及pii、privacy和french等标签。数据集规模在10K到100K之间,具体包括训练集(57,248个例子)、验证集(500个例子)和测试集(2,500个例子),以及一个英语回归检查集(426个例子)。数据集定义了8个PII类别,每个类别在测试集中至少有100个跨度。数据格式为OPF JSONL,兼容`opf eval`。数据来源于三个不同的数据集,经过去重和分层处理,确保测试集的严格独立性。
The French PII Evaluation Dataset is a curated dataset for evaluating and training French Personally Identifiable Information (PII) detection, specifically designed for benchmarking the TheoDB/privacy-filter-fr model. The dataset includes content in both French and English, targeting token-classification tasks with tags such as pii, privacy, and french. The dataset size ranges between 10K and 100K, comprising a training set (57,248 examples), a validation set (500 examples), and a test set (2,500 examples), along with an English regression check set (426 examples). It defines 8 PII classes, each with at least 100 spans in the test set. The data is formatted in OPF JSONL, compatible with `opf eval`. The dataset sources include three different datasets, which were deduplicated and stratified to ensure the test sets strict independence.



