RAT-Bench
收藏资源简介:
RAT-Bench是由伦敦帝国理工学院开发的一个综合性文本匿名化评估基准数据集。该数据集基于美国人口普查局的1%公共使用微数据样本(PUMS)生成,包含合成文本,涵盖不同领域、语言和难度级别的直接与间接标识符。数据集通过模拟真实人口统计分布,支持对匿名化工具的重识别风险进行量化评估,尤其关注法律合规性要求的k-匿名性(k=5)。其应用领域聚焦于隐私保护技术研发,旨在解决现有匿名化工具对非标准表述标识符识别不足、跨语言泛化能力弱等问题,为AI模型训练数据脱敏提供标准化测试环境。
RAT-Bench is a comprehensive text anonymization evaluation benchmark dataset developed by Imperial College London. Built upon the 1% Public Use Microdata Sample (PUMS) from the U.S. Census Bureau, the dataset contains synthetic texts covering direct and indirect identifiers across diverse domains, languages and difficulty levels. By simulating real demographic distributions, it enables quantitative assessment of re-identification risks for anonymization tools, with a particular focus on k-anonymity (k=5) required by legal compliance standards. Targeted at privacy protection technology research and development, RAT-Bench aims to address the shortcomings of existing anonymization tools, such as insufficient recognition of non-standardly expressed identifiers and poor cross-language generalization capabilities, and provides a standardized testing environment for data anonymization in AI model training.



