AmazonScience/ESRRSim
收藏资源简介:
ESRRSim生成基准是一个用于评估大型语言模型中新兴战略推理风险(ESRRs)的数据集,由ESRRSim代理框架生成。该数据集包含1,052个评估场景,每个场景都配有双评分标准,用于评估模型响应和推理轨迹。数据集设计为法官无关的,评分标准指定了具体的行为准则,适用于任何LLM法官或人类评估者。它涵盖了7个风险类别(如奖励黑客、欺骗、评估游戏等)和20个子类别,以及6个场景类型(如欺骗需求游戏、伦理困境等)。每个记录以JSON格式存储,包括评估提示、模型响应评分标准、思考响应评分标准和元数据。数据集主要用于评估LLM的行为风险模式,所有内容均为虚构和合成生成。
The ESRRSim Generated Benchmark is an evaluation benchmark for Emergent Strategic Reasoning Risks (ESRRs) in large language models, generated by the ESRRSim agentic framework. This dataset provides 1,052 evaluation scenarios with paired dual rubrics for assessing both model responses and reasoning traces. It is judge-agnostic, with rubrics specifying concrete behavioral criteria applicable by any LLM judge or human evaluator. The benchmark covers 7 risk categories (e.g., Reward Hacking, Deception, Evaluation Gaming) with 20 subcategories, and 6 scenario types (e.g., Deception-Required Games, Ethical Dilemmas). Each record is in JSON format, including evaluation prompt, model response rubric, thought response rubric, and metadata. It is designed exclusively for evaluating behavioral risk patterns in LLMs, with all content being fictional and synthetically generated.




