MMDiag
收藏资源简介:
MMDiag是一个多轮多模态对话数据集,由北京大学计算机科学与技术学院和北京人工智能科学院合作创建。该数据集包含日常场景、表格场景和Minigrid场景三种类型,旨在测试多模态大语言模型在处理具有挑战性的多轮对话时的推理能力。数据集通过规则搜索和GPT-4o-mini的辅助生成,具有问题之间、问题与图像之间以及不同图像区域之间的强相关性,更贴近现实世界的场景。
MMDiag is a multi-turn multimodal dialogue dataset jointly developed by the School of Computer Science and Technology, Peking University and the Beijing Academy of Artificial Intelligence. It encompasses three scenario types: daily scenarios, tabular scenarios, and Minigrid scenarios, designed to test the reasoning capabilities of multimodal large language models when processing challenging multi-turn dialogues. The dataset is created through rule-based searches and assisted generation with GPT-4o-mini, and it features strong correlations between different questions, between questions and their corresponding images, as well as among distinct image regions, making it more aligned with real-world scenarios.

- 1Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning北京大学计算机科学与技术学院, 北京人工智能科学院 · 2025年



