M2RAG
收藏资源简介:
M2RAG数据集由北京理工大学计算机科学与技术学院创建,旨在评估多模态生成模型的能力。该数据集包含200个查询样本,涵盖11个不同类别,每个查询样本都附有相关的多模态网页和辅助图像。数据集的创建过程包括查询收集、数据准备和元素评估三个步骤,确保了数据的高质量和多样性。M2RAG数据集主要应用于多模态生成任务,旨在通过结合文本和图像信息,提升生成内容的信息密度和可读性。
The M2RAG dataset was constructed by the School of Computer Science and Technology, Beijing Institute of Technology, with the primary goal of evaluating the capabilities of multimodal generative models. This dataset comprises 200 query samples spanning 11 distinct categories, with each query sample paired with relevant multimodal webpages and auxiliary images. The dataset creation process includes three core steps: query collection, data preparation, and element evaluation, which ensures the high quality and diversity of the dataset. The M2RAG dataset is primarily utilized for multimodal generative tasks, aiming to enhance the information density and readability of generated content by combining textual and visual information.




