AMBENCH
收藏资源简介:
AMBENCH是一个由看似模棱两可的人名组成的基准数据集,这些名字利用了名字常规性偏差现象,嵌入在简短文本片段中,并伴有良性提示注入。该数据集旨在评估大型语言模型在检测个人身份信息方面的能力,特别是在模糊上下文中。实验表明,现代大型语言模型在识别模棱两可的人名时,召回率比识别更易辨认的名字低20-40%。此外,当良性提示注入存在时,模棱两可的人名在LLM生成的隐私保护摘要中被忽略的可能性是其他名字的四倍。这些发现突显了完全依赖LLM来保护用户隐私的风险,并强调了需要对其隐私失败模式进行更系统的研究。
AMBENCH is a benchmark dataset composed of seemingly ambiguous personal names. These names leverage the phenomenon of nominal regularity bias, embedded within short text snippets alongside benign prompt injections. This dataset is designed to evaluate the ability of large language models (LLMs) to detect personally identifiable information (PII), particularly in ambiguous contexts. Experiments demonstrate that modern LLMs have a 20-40% lower recall rate when identifying ambiguous personal names compared to more readily distinguishable ones. Additionally, when benign prompt injections are present, ambiguous personal names are four times more likely to be omitted from privacy-preserving summaries generated by LLMs than other names. These findings highlight the risks of fully relying on LLMs for protecting user privacy, and underscore the need for more systematic research into their privacy failure modes.
数据集概述
基本信息
- 数据集名称: Can Large Language Models Really Recognize Your Name?
- 数据集地址: https://github.com/dzungvpham/llm-name-detection
- 论文地址: https://arxiv.org/abs/2505.14549
数据集状态
- 当前状态: 代码和数据即将添加(Code and data will be added soon!)
相关研究
- 研究主题: 大型语言模型是否能真正识别姓名

- 1Can Large Language Models Really Recognize Your Name?Google Research · 2025年



