MonitorBench
收藏资源简介:
MonitorBench是由伊利诺伊大学厄巴纳-香槟分校等机构联合推出的首个全开源、综合性基准测试数据集,旨在系统评估大型语言模型中思维链(CoT)的可监控性。该数据集包含1,514条测试实例,涵盖19个任务和7个类别,通过精心设计的决策关键因素来刻画CoT监控模型行为的适用场景。数据集构建过程包括标准测试和压力测试两种设置,后者用于量化CoT可监控性的退化程度。该数据集主要应用于自然语言处理领域,解决LLM推理过程中思维链与最终输出因果脱节导致的监控失效问题,为开发新型监控方法提供研究基础。
MonitorBench is the first fully open-source, comprehensive benchmark dataset jointly launched by the University of Illinois Urbana-Champaign and other institutions, aiming to systematically evaluate the monitorability of Chain-of-Thought (CoT) in Large Language Models (LLMs). This dataset contains 1,514 test instances covering 19 tasks and 7 categories, and characterizes applicable scenarios for monitoring CoT model behaviors via carefully designed decision-critical factors. The dataset construction includes two testing settings: standard test and stress test, where the latter is used to quantify the degree of degradation in CoT monitorability. This dataset is primarily applied in the field of natural language processing, addressing the monitoring failure problem caused by the causal disconnect between Chain-of-Thought and final outputs during LLM reasoning, and providing a research foundation for developing novel monitoring methods.
MonitorBench 数据集概述
数据集基本信息
- 数据集名称:MonitorBench
- 核心目标:评估大型语言模型中思维链的可监控性
- 论文地址:https://arxiv.org/pdf/2603.28590
- 许可协议:Apache 2.0
数据集内容与构成
- 测试实例数量:1,514个
- 任务覆盖:涵盖7个类别下的19个任务
- 关键设计:包含精心设计的决策关键因素,用于表征何时可以利用思维链来监控驱动LLM行为的因素
- 压力测试:包含两种压力测试设置,用于量化思维链可监控性在多大程度上可能被削弱
主要功能与用途
- 提供一个多样化的测试集,用于评估思维链在监控LLM行为驱动因素方面的适用性。
- 通过压力测试量化思维链监控能力的潜在退化程度。
当前状态
- 论文已发布。
- 代码、基准测试实例及环境安装脚本计划于四月发布。
- 支持自定义数据集的说明将在后续准备。
作者与机构
- 主要贡献者(同等贡献):Han Wang, Yifan Sun, Brian Ko
- 其他作者:Mann Talati, Jiawen Gong, Zimeng Li, Naicheng Yu, Xucheng Yu, Wei Shen, Vedant Jolly, Huan Zhang
- 参与机构:伊利诺伊大学厄巴纳-香槟分校、华盛顿大学、加州大学圣地亚哥分校
联系与引用
- 联系邮箱:hanw14@illinois.edu
- 引用格式:请使用提供的BibTeX条目进行引用。




