MTRAG
收藏资源简介:
MTRAG是由IBM研究院开发的多轮对话检索增强生成(RAG)基准数据集,旨在评估RAG系统在多轮对话中的表现。该数据集包含110个对话,平均每个对话有7.7轮,总共842个任务,涵盖了四个不同领域(如维基百科、金融、政府和技术文档)。数据集的创建过程通过人工标注者与RAG系统的实时交互完成,确保了对话的多样性和真实性。每个对话都经过精心设计,包含多种问题类型、多轮对话模式以及可回答性维度。MTRAG的应用领域主要集中在自然语言处理中的对话系统评估,旨在解决多轮对话中检索和生成的挑战,特别是在处理不可回答问题、非独立问题以及跨领域对话时的表现。
MTRAG is a multi-turn dialogue retrieval-augmented generation (RAG) benchmark dataset developed by IBM Research, designed to evaluate the performance of RAG systems in multi-turn dialogue scenarios. This dataset includes 110 dialogues, with an average of 7.7 turns per dialogue and a total of 842 tasks, covering four distinct domains such as Wikipedia, finance, government, and technical documentation. The dataset was constructed through real-time interaction between human annotators and RAG systems, ensuring the diversity and authenticity of the dialogues. Each dialogue is meticulously designed to incorporate multiple question types, multi-turn dialogue patterns, and answerability dimensions. The main application scope of MTRAG focuses on the evaluation of dialogue systems in natural language processing, aiming to address the challenges of retrieval and generation in multi-turn dialogues, especially the performance when handling unanswerable questions, non-independent questions, and cross-domain dialogues.

- 1MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation SystemsIBM研究院 · 2025年



