aisa-group/EvalAwareBench
收藏资源简介:
EvalAwareBench是一个因素控制的基准测试数据集,用于研究语言模型中的评估意识。该数据集基于心理学基础,设计了八个可独立切换的触发因素,以测量模型对评估线索的识别、行为一致性以及这些线索如何组合影响模型响应。数据集包含100个配对任务(每个任务有安全变体和能力变体,共200个任务模板),每个任务有8个独立可控因素(F1-F8),每个任务变体有256个因素配置(2^8组合),总共有51,200个渲染提示。数据集的目的是系统性地分析评估意识,从单因素消融到全组合分析,帮助理解语言模型在遇到不同评估线索时的行为变化。
A factor-controlled benchmark for studying evaluation awareness in language models, where eight psychology-grounded trigger factors can be independently toggled on matched safety and capability tasks to measure recognition, behavioral consistency, and how evaluative cues combine. The dataset includes 100 paired tasks (safety + capability variants = 200 task templates), 8 independently controllable factors (F1–F8) per task, 256 factor configurations per task variant (2^8 combinations), and 51,200 total rendered prompts across all tasks and configurations.



