REAL-MM-RAG
收藏资源简介:
REAL-MM-RAG数据集是由IBM Research Israel和Weizmann Institute of Science创建的多模态检索基准。该数据集包含8000页文档,涵盖四个子领域,旨在满足真实世界检索的四个关键特性:多模态文档、增强难度、真实RAG查询和准确标注。数据集包括文本、图表、表格和图像等,要求系统处理组合的文本和视觉数据。数据集通过自动化管道进行查询生成、过滤、重写和错误验证,以提供可靠的评估和多级别的查询重写鲁棒性评估。
The REAL-MM-RAG dataset is a multimodal retrieval benchmark developed by IBM Research Israel and the Weizmann Institute of Science. This dataset consists of 8,000 pages of documents across four subfields, and is designed to meet four key characteristics of real-world retrieval: multimodal documents, enhanced difficulty, realistic RAG queries, and accurate annotations. It includes text, charts, tables, images and other modalities, requiring the system to process combined textual and visual data. The dataset adopts automated pipelines for query generation, filtering, rewriting and error validation to deliver reliable evaluation and robust assessment of multi-level query rewriting.




