TIGER-Lab/Mantis-Eval
收藏资源简介:
Mantis-Eval是一个新策划的数据集,用于评估多模态语言模型在多图像推理上的能力。该数据集包含200多个由人类注释的挑战性多图像推理问题。数据集的特征包括id、问题类型、问题、图像、选项、答案、数据来源和类别。数据集的分割信息显示,测试集包含217个示例,总字节数为479770102。
Mantis-Eval is a newly curated dataset to evaluate multimodal language models capability to reason over multiple images. This evaluation dataset contains more than 200 human-annotated challenging multi-image reasoning problems. The features of the dataset include id, question type, question, images, options, answer, data source, and category. The split information shows that the test set contains 217 examples with a total of 479770102 bytes.
数据集概述
基本信息
- 语言: 英语
- 许可证: Apache 2.0
- 大小分类: n<1K
- 任务分类: 问答
- 美观名称: Mantis-Eval
数据集配置
- 配置名称: mantis_eval
- 特征:
- id: 字符串
- question_type: 字符串
- question: 字符串
- images: 图像序列
- options: 字符串序列
- answer: 字符串
- data_source: 字符串
- category: 字符串
- 分割:
- test:
- 字节数: 479770102
- 示例数: 217
- test:
- 下载大小: 473031413
- 数据集大小: 479770102
数据文件
- 配置名称: mantis_eval
- 数据文件:
- 分割: test
- 路径: mantis_eval/test-*
统计信息
- 包含超过200个人工标注的复杂多图像推理问题。
排行榜
| 模型 | 大小 | Mantis-Eval |
|---|---|---|
| GPT-4V | - | 62.67 |
| Mantis-SigLIP | 8B | 59.45 |
| Mantis-Idefics2 | 8B | 57.14 |
| Mantis-CLIP | 8B | 55.76 |
| VILA | 8B | 51.15 |
| BLIP-2 | 13B | 49.77 |
| Idefics2 | 8B | 48.85 |
| InstructBLIP | 13B | 45.62 |
| LLaVA-V1.6 | 7B | 45.62 |
| CogVLM | 17B | 45.16 |
| Qwen-VL-Chat | 7B | 39.17 |
| Emu2-Chat | 37B | 37.79 |
| VideoLLaVA | 7B | 35.04 |
| Mantis-Flamingo | 9B | 32.72 |
| LLaVA-v1.5 | 7B | 31.34 |
| Kosmos2 | 1.6B | 30.41 |
| Idefics1 | 9B | 28.11 |
| Fuyu | 8B | 27.19 |
| OpenFlamingo | 9B | 12.44 |
| Otter-Image | 9B | 14.29 |
引用
如果使用此数据集,请引用以下工作:
@inproceedings{Jiang2024MANTISIM, title={MANTIS: Interleaved Multi-Image Instruction Tuning}, author={Dongfu Jiang and Xuan He and Huaye Zeng and Cong Wei and Max W.F. Ku and Qian Liu and Wenhu Chen}, publisher={arXiv2405.01483} year={2024}, }




