遇见数据集

aisa-group/EvalAwareBench

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

EvalAwareBench是一个因素控制的基准测试数据集,用于研究语言模型中的评估意识。该数据集基于心理学基础,设计了八个可独立切换的触发因素,以测量模型对评估线索的识别、行为一致性以及这些线索如何组合影响模型响应。数据集包含100个配对任务(每个任务有安全变体和能力变体,共200个任务模板),每个任务有8个独立可控因素(F1-F8),每个任务变体有256个因素配置(2^8组合),总共有51,200个渲染提示。数据集的目的是系统性地分析评估意识,从单因素消融到全组合分析,帮助理解语言模型在遇到不同评估线索时的行为变化。

A factor-controlled benchmark for studying evaluation awareness in language models, where eight psychology-grounded trigger factors can be independently toggled on matched safety and capability tasks to measure recognition, behavioral consistency, and how evaluative cues combine. The dataset includes 100 paired tasks (safety + capability variants = 200 task templates), 8 independently controllable factors (F1–F8) per task, 256 factor configurations per task variant (2^8 combinations), and 51,200 total rendered prompts across all tasks and configurations.

提供机构:
aisa-group
二维码
社区交流群
二维码
科研交流群
商业服务