mmRAG
收藏资源简介:
mmRAG是一个模块化的基准数据集,旨在评估多模态检索增强生成(RAG)系统。该数据集集成了来自六个不同问答数据集的查询,涵盖了文本、表格和知识图谱。数据集包含5124548条数据,包括5,124个查询、3.2百万个文档块和88,751个已标注的查询-块对。mmRAG的构建过程包括数据集收集、数据加工和数据标注三个阶段。数据集的标注采用标准的信息检索协议,提供了充分的标注信息来评估检索和查询路由的准确性。该数据集可用于评估RAG系统的查询路由、检索和生成等主要组件,为多模态RAG系统的模块化评估提供了一个独特的测试平台。
mmRAG is a modular benchmark dataset designed to evaluate multimodal retrieval-augmented generation (RAG) systems. It integrates queries from six distinct question answering datasets, covering text, tables, and knowledge graphs. The dataset comprises 5,124,548 total entries, including 5,124 queries, 3.2 million document chunks, and 88,751 annotated query-chunk pairs. The construction of mmRAG includes three stages: dataset collection, data processing, and data annotation. The dataset adopts standard information retrieval protocols for annotation, providing sufficient annotated information to evaluate the accuracy of retrieval and query routing. This dataset can be used to evaluate core components of RAG systems such as query routing, retrieval, and generation, providing a unique testbed for modular evaluation of multimodal RAG systems.




