NLPCoreTeam/ruMT-Bench
收藏资源简介:
ruMT-Bench包含8个不同知识领域(写作、角色扮演、提取、推理、数学、编码、STEM、人文/社会科学)的多轮指导性问题。GPT-4对模型的回答进行1到10的评分,最终得分由整个对话的平均分决定。对于一些需要精确答案的复杂问题(如数学和编码),评分提示中包含参考答案以帮助评估大语言模型的回答。数据集的局限性包括对长回答的偏好、自我增强偏见、在数学和推理问题评分上的限制以及每个类别问题数量的限制。
ruMT-Bench contains multi-turn instructional questions across 8 distinct knowledge domains: writing, role-playing, information extraction, reasoning, mathematics, coding, STEM, and humanities/social sciences. GPT-4 rates the model's responses on a scale of 1 to 10, and the final score is determined by the average score across the entire conversation. For some complex questions requiring precise answers such as those in mathematics and coding, reference answers are included in the scoring prompt to assist in evaluating the responses of large language models. The limitations of the dataset include a preference for long responses, self-enhancement bias, constraints on scoring for mathematical and reasoning questions, and a limited number of questions per category.
ruMT-Bench 数据集概述
基本信息
- 许可证: Apache-2.0
- 任务类别: 问答
- 语言: 俄语
- 标签: 评估
- 美观名称: ruMT-Bench
- 大小类别: 小于1K
数据集内容
- 内容描述: 包含8个不同知识领域的多轮问题,包括写作、角色扮演、提取、推理、数学、编程、STEM、人文/社会科学。
- 评分机制: 使用GPT-4对模型响应进行1到10分的评分,最终得分是整个对话的平均分。对于需要精确答案的复杂问题(如数学和编程),提供参考答案以辅助评估。
数据集配置
- 配置名称: default
- 数据文件:
- 分割: 测试
- 路径: question.jsonl
数据集局限性
- 冗长偏差: LLM评估者偏好较长答案,即使它们不如短答案好。GPT-4在处理长度偏差方面表现更好。
- 自我增强偏差: GPT-4在自我评分时胜率更高,而Claude偏好自身答案25%,GPT-3.5则不偏好自身答案。
- 评估能力限制: 在评估数学和推理问题时能力有限,评估质量受评估者能力限制。
- 样本量限制: 每个类别仅包含10个问题(20个问题),可能无法全面代表所有LLM能力。




