MMDU
收藏资源简介:
MMDU是一个专为评估和提升大型视觉语言模型(LVLMs)在多轮多图像对话理解能力而设计的综合基准和大规模指令调优数据集。该数据集由上海人工智能实验室创建,包含45,000条高质量数据,旨在通过模拟真实世界的人机交互场景,测试和改进模型在处理多图像和长对话历史中的表现。数据集通过使用聚类算法从开放源代码的维基百科中提取相关图像和文本描述,并由人工注释者借助GPT-4o模型构建问答对。MMDU不仅挑战了现有LVLMs的处理能力,还通过其开放式的评估方式,推动了模型在理解和生成自然、有意义对话方面的进步。
MMDU is a comprehensive benchmark and large-scale instruction-tuning dataset specifically designed to evaluate and enhance the multi-turn, multi-image dialogue understanding capabilities of Large Vision-Language Models (LVLMs). Developed by the Shanghai AI Laboratory, this dataset contains 45,000 high-quality samples. It aims to test and improve models' performance in handling multiple images and long conversation histories by simulating real-world human-computer interaction scenarios. The dataset extracts relevant images and textual descriptions from open-source Wikipedia using clustering algorithms, and human annotators construct question-answer pairs with the assistance of GPT-4o. MMDU not only challenges the processing capabilities of existing LVLMs but also promotes the advancement of models in understanding and generating natural, meaningful dialogues through its open-ended evaluation approach.

- 1MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs上海人工智能实验室 · 2024年



