Comprehensive Arabic Multimodal Reasoning Benchmark (ARB)
收藏资源简介:
ARB是一个全面的阿拉伯语多模态推理基准数据集,旨在评估阿拉伯语中多模态推理的逐步推理过程。该数据集涵盖了11个不同的领域,包括视觉推理、文档理解、OCR、科学分析和文化解释。ARB包含1,356个多模态样本,配对5,119个人工编辑的推理步骤和相应的动作。该数据集提供了一个结构化的框架,用于诊断在代表性不足的语言中进行多模态推理,并标志着迈向包容性、透明性和文化意识的人工智能系统的重要一步。
ARB is a comprehensive Arabic multimodal reasoning benchmark dataset designed to assess the step-by-step reasoning procedures for multimodal reasoning tasks in Arabic. This dataset spans 11 distinct domains, including visual reasoning, document understanding, OCR, scientific analysis, and cultural interpretation. ARB contains 1,356 multimodal samples, which are paired with a total of 5,119 manually edited reasoning steps and their corresponding actions. This dataset provides a structured framework for diagnosing multimodal reasoning in underrepresented languages, and marks a significant step toward inclusive, transparent, and culturally aware AI systems.
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
概述
- 名称: ARB (A Comprehensive Arabic Multimodal Reasoning Benchmark)
- 类型: 多模态推理基准
- 语言: 阿拉伯语
- 模态: 文本和视觉
- 目标: 评估阿拉伯语多模态模型的逐步推理能力
- 特点: 首个针对阿拉伯语多模态逐步推理的基准,涵盖11个不同领域
数据集详情
- 样本数量: 1,356个多模态样本
- 推理步骤: 5,119个精心策划的推理步骤
- 领域覆盖: 11个不同领域,包括:
- 视觉推理
- OCR和文档理解
- 图表和图解解释
- 数学和逻辑推理
- 科学和医学分析
- 文化和历史解释
- 遥感
- 农业图像分析
- 复杂视觉感知
数据分布
- 数学与逻辑: 41%
- 图表、图解与表格: 24%
- 其他领域: 包括社会与文化、科学、医学等
数据来源
- 英语推理基准
- 阿拉伯语问答基准
- 英语字幕数据集
- 合成数据
- 工具增强数据
评估指标
- 核心维度:
- 忠实度 (At-Tat¯abuq)
- 信息量 (Al-Ithr¯a’ Al-Ma’l¯um¯at¯ı)
- 连贯性 (At-Taw¯afuq)
- 常识 (Al-Mantiq Al-’A¯mm)
- 推理对齐 (At-Tawa¯fuq Al-Istidla¯l¯ı)
- 辅助检查:
- 幻觉
- 冗余
- 语义差距
- 缺失步骤
评估结果
闭源模型
| 模型 | 最终答案准确率 (%) | 推理步骤质量 (%) |
|---|---|---|
| GPT-4o | 60.22 | 64.29 |
| GPT-4o-min | 52.22 | 61.02 |
| GPT-4.1 | 59.43 | 80.41 |
| o4-mini | 58.93 | 80.75 |
| Gemini 1.5 Pro | 56.70 | 64.34 |
| Gemini 2.0 Flash | 57.80 | 64.09 |
开源模型
| 模型 | 最终答案准确率 (%) | 推理步骤质量 (%) |
|---|---|---|
| Qwen2.5VL-7b | 37.02 | 64.03 |
| Llama-3.2-11B-Vis-Inst. | 25.58 | 53.20 |
| AIN | 27.35 | 52.77 |
| Llama-4-Scout-17Bx16E | 48.52 | 77.70 |
| Aya-Vision-8B | 28.81 | 63.64 |
| InternVl3-8B | 31.04 | 54.50 |
引用
bibtex @misc{ghaboura2025arbcomprehensivearabicmultimodal, title={ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark}, author={Sara Ghaboura and Ketan More and Wafa Alghallabi and Omkar Thawakar and Jorma Laaksonen and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer}, year={2025}, eprint={2505.17021}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2505.17021}, }



