SCPI Dataset
收藏资源简介:
SCPI数据集是由早稻田大学与国立信息学研究所联合构建的、专门针对日本《个人信息保护法》所定义的特殊需谨慎个人信息进行标注的文本数据集。该数据集源自日本Common Crawl网络语料库的采样,包含约261万条文本,平均每条长度为2474字符,最终通过两阶段大语言模型标注流程筛选出涵盖医疗史、犯罪记录、残疾状况等核心敏感类别的954条正例样本,并配以8586条负例构成平衡训练集。数据集的创建过程首先基于命名实体识别模型筛选含人名文本,随后采用低参数模型与高性能模型协同的标注策略,并辅以人工校正确保标注质量。该数据集主要应用于大型语言模型预训练语料库的隐私合规过滤,旨在开发高效SCPI分类器以自动检测并移除日语文本中的敏感个人信息,从而防范模型记忆泄露风险并满足日本法律监管要求。
The SCPI dataset is a text annotation dataset jointly constructed by Waseda University and the National Institute of Informatics, specifically targeting special care-required personal information as defined by Japan’s Act on the Protection of Personal Information. This dataset is sampled from the Japanese Common Crawl web corpus, containing approximately 2.61 million text entries with an average length of 2474 characters per entry. Finally, 954 positive samples covering core sensitive categories including medical history, criminal records, and disability status were screened out through a two-stage large language model annotation pipeline, paired with 8586 negative samples to form a balanced training dataset. The dataset construction process first filters text containing personal names using a Named Entity Recognition (NER) model, then adopts a collaborative annotation strategy combining low-parameter and high-performance models, supplemented by manual verification to ensure annotation quality. This dataset is mainly applied to privacy compliance filtering for large language model pre-training corpora, aiming to develop an efficient SCPI classifier to automatically detect and remove sensitive personal information in Japanese texts, thereby preventing the risk of model memorization leakage and complying with Japanese legal and regulatory requirements.




