RefuteBench 2.0
收藏资源简介:
RefuteBench 2.0是一个由浙江大学和西湖大学联合创建的动态评估框架,用于评估大型语言模型(LLM)对反驳指令的遵循能力。该数据集包含了机器翻译、总结和开放式写作等多种任务类型,通过代理生成的反驳和评估,模拟真实世界的多轮对话场景,以评估LLM在处理用户反馈时的性能和适应性。数据集涵盖了300个种子问题,以及对应的对话轮次和 tokens 数,旨在解决LLM在长时间对话中保留和使用先前信息的能力问题。
RefuteBench 2.0 is a dynamic evaluation framework jointly developed by Zhejiang University and Westlake University, designed to assess the ability of large language models (LLMs) to follow rebuttal instructions. This dataset covers multiple task types including machine translation, text summarization, and open-ended writing. It simulates real-world multi-turn dialogue scenarios via agent-generated rebuttals and evaluations, to gauge the performance and adaptability of LLMs when handling user feedback. The dataset includes 300 seed questions, along with their corresponding dialogue turns and token counts, aiming to address the issue of LLMs' capacity to retain and utilize prior information during prolonged conversations.




