REVERSO
收藏资源简介:
REVERSO是一个新颖的基准数据集,用于评估大型语言模型是否能够遵循反事实指令来模拟具有反转表现的角色。该数据集以数学推理为代表性场景,改编自广泛采用的GSM8k数据集。REVERSO旨在通过模拟在数学推理方面具有高表现和低表现的学生,来评估LLM的反事实指令遵循能力。数据集还包括一个交叉设置,其中模型需要额外模拟角色的种族背景,以测试这种背景是否会影响模型模拟反转表现角色的能力。
REVERSO is a novel benchmark dataset developed to evaluate whether large language models (LLMs) can follow counterfactual instructions to simulate characters with reversed performance. Adopting mathematical reasoning as its representative scenario, the dataset is adapted from the widely adopted GSM8K benchmark. REVERSO aims to assess the counterfactual instruction-following capability of LLMs by simulating students with either high or low performance in mathematical reasoning. Additionally, the dataset includes a cross setting where models are required to further simulate the racial backgrounds of the characters, to test whether such backgrounds will affect the model's ability to simulate characters with reversed performance.




