MIRAGE-BENCH
收藏资源简介:
MIRAGE-BENCH是由滑铁卢大学和Vectara共同创建的多语言检索增强生成(RAG)系统基准测试数据集,涵盖18种不同语言。数据集基于MIRACL检索数据集构建,包含11195条评估数据和39763条训练数据。数据集的创建过程包括使用MIRACL中的查询和相关性判断,并通过GPT-4o等大型语言模型生成多语言答案。MIRAGE-BENCH主要用于评估多语言RAG系统在生成任务中的表现,旨在解决现有RAG基准测试主要集中在英语上的问题。
MIRAGE-BENCH is a multilingual Retrieval-Augmented Generation (RAG) system benchmark dataset co-developed by the University of Waterloo and Vectara, spanning 18 distinct languages. Constructed on the basis of the MIRACL retrieval dataset, it consists of 11,195 evaluation instances and 39,763 training instances. The dataset creation workflow employs queries and relevance judgments sourced from MIRACL, and generates multilingual answers using large language models (LLMs) such as GPT-4o. MIRAGE-BENCH is primarily designed to assess the performance of multilingual RAG systems in generation tasks, with the goal of addressing the limitation that existing RAG benchmarks are predominantly focused on the English language.
MIRAGE-BENCH
概述
- 名称: MIRAGE-BENCH
- 描述: 用于论文《MIRAGE-BENCH: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems》的代码和数据集。
状态
- 发布状态: 代码和数据即将发布。

- 1MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems滑铁卢大学, 加拿大; Vectara, 美国 · 2024年



