MemTrapBench
收藏资源简介:
MemTrapBench是浙江大学等机构联合构建的评估大语言模型记忆认知陷阱的基准数据集,包含1050个精心设计的实例,涵盖推理固化和信念扭曲两大类别下的四个场景:认知偏差、创伤、任务边界与安全。数据集通过种子设计、多轮对抗性对话生成及两阶段质量控制(自动过滤与专家审核)构建而成,每个实例均标注有标准答案与预期失败模式。该基准旨在系统评估记忆如何扭曲模型推理或信念,从而降低当前任务性能,揭示现有记忆框架在认知陷阱面前的脆弱性。
MemTrapBench is a benchmark dataset for evaluating memory-induced cognitive traps in large language models (LLMs), jointly constructed by Zhejiang University and other institutions. It includes 1050 meticulously crafted instances spanning four scenarios under two major categories: reasoning fixation and belief distortion, specifically cognitive biases, trauma, task boundaries, and safety. The dataset is constructed via seed design, multi-round adversarial dialogue generation, and a two-stage quality control pipeline consisting of automatic filtering and expert review. Each instance is annotated with a standard answer and an expected failure mode. This benchmark aims to systematically evaluate how memory distorts model reasoning or beliefs, thereby compromising current task performance, and uncover the vulnerability of existing memory frameworks when faced with cognitive traps.
MemTrapBench 数据集概述
数据集名称:MemTrapBench
所属机构:浙江大学 NLP 实验室(ZJUNLP)
数据集地址:https://github.com/zjunlp/MemTrapBench
核心主题:大语言模型(LLM)记忆使用中的认知陷阱基准测试
1. 数据集目标
- 旨在系统性地评估和衡量大语言模型在记忆使用过程中可能出现的认知陷阱(Cognitive Traps)。
- 通过构建基准测试,促进对 LLM 记忆机制的深入理解,并暴露模型在记忆调用、推理和决策中的潜在缺陷。
2. 适用场景
- 大语言模型的性能评估与鲁棒性测试。
- 认知科学、人工智能安全、记忆机制研究。
- 模型训练与调优过程中的诊断工具。
3. 数据集内容
- 该基准测试专注于“记忆使用”这一特定认知环节,重点覆盖可能引发错误或偏差的陷阱场景。
- 具体数据样例、任务形式及评估指标在仓库中尚未详细展示(README 仅提供名称与一句描述),需进一步查看仓库文件以获取完整信息。
4. 获取与使用
- 数据及代码通过 GitHub 仓库开放,可直接访问上述地址获取全部资源。
- 仓库结构、许可协议、引用方式等详细信息需在仓库页面中查看。




