ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
收藏资源简介:
ARB是第一个专注于阿拉伯语跨文本和视觉模态逐步推理的基准,涵盖科学、文化、OCR和历史解释等11个不同领域。它包括1,356个多模态样本,每个样本包含一个图像、阿拉伯语问题和基于推理的答案,以及5,119个精心策划的推理步骤。数据集由阿拉伯语母语者和领域专家验证,包含原始阿拉伯语数据、高质量翻译和合成样本。
ARB is the first benchmark dedicated to cross-text and visual modality incremental reasoning in Arabic, encompassing 11 diverse domains including science, culture, OCR, and historical interpretation. It includes 1,356 multimodal samples, each with an image, an Arabic question, and an answer based on reasoning, as well as 5,119 meticulously crafted reasoning steps. The dataset has been verified by Arabic native speakers and domain experts, and it contains original Arabic data, high-quality translations, and synthetic samples.
ARB: 阿拉伯多模态推理基准数据集
数据集概述
- 名称: ARB (A Comprehensive Arabic Multimodal Reasoning Benchmark)
- 类型: 多模态基准数据集
- 语言: 阿拉伯语
- 模态: 文本和视觉
- 样本数量: 1,356个多模态样本
- 推理步骤: 5,119个
关键特性
- 多样性: 覆盖11个不同领域,包括科学、文化、OCR和历史解释等
- 验证: 由阿拉伯语母语者和领域专家验证
- 数据来源: 原始阿拉伯语数据、高质量翻译和合成样本的混合
- 开放性: 完全开源的数据集和工具包
数据构成
- 每个样本包含:
- 图像
- 阿拉伯语问题
- 基于推理的答案
- 选项(针对多选题)
- 有序推理链
- 最终解决方案(阿拉伯语)
- 领域类别(11个类别之一)
- 课程类型(4种之一)
领域分布
| 领域 | 英文基准 | 阿拉伯基准 | 人工创建 | 合成 |
|---|---|---|---|---|
| 视觉推理 | ✅ | – | – | – |
| OCR与文档分析 | – | – | ✅ | ✅ |
| 图表与数据表(CDT) | ✅ | ✅ | ✅ | ✅ |
| 数学与逻辑 | ✅ | – | – | – |
| 社会与文化 | ✅ | – | – | – |
| 计算机视觉感知 | ✅ | – | – | – |
| 医学图像分析 | ✅ | ✅ | – | – |
| 科学推理 | ✅ | – | – | – |
| 农业解释 | ✅ | – | ✅ | ✅ |
| 遥感理解 | – | ✅ | – | – |
| 历史与人类学 | ✅ | – | ✅ | ✅ |
评估协议
- 评估指标:
- 词法和语义相似度分数(BLEU、ROUGE、BERTScore)
- 跨语言语义对齐(LaBSE)
- 自定义阿拉伯语评估标准(包括10个因素)
- LLM评估:
- 逐步推理质量(连贯性、信息量、常识)
- 最终答案准确性
- 与人类评分者的一致性(Krippendorffs Alpha > 87%)
评估结果
闭源模型
| GPT-4o | GPT-4o-mini | GPT-4.1 | o4-mini | Gemini 1.5 Pro | Gemini 2.0 Flash | |
|---|---|---|---|---|---|---|
| 最终答案 (%) | 60.22 | 52.22 | 59.43 | 58.93 | 56.7 | 57.8 |
| 推理步骤 (%) | 64.29 | 61.02 | 80.41 | 80.75 | 64.34 | 64.09 |
开源模型
| Qwen2.5-VL-7B | Llama-3.2-11B | AIN | Llama-4 Scout | Aya-Vision-8B | InternVL3-8B | |
|---|---|---|---|---|---|---|
| 最终答案 (%) | 37.02 | 25.58 | 27.35 | 48.52 | 28.81 | 31.04 |
| 推理步骤 (%) | 64.03 | 53.2 | 52.77 | 77.7 | 63.64 | 54.5 |
下载
bash from datasets import load_dataset ds = load_dataset("MBZUAI/ARB")
引用
bibtex @misc{ghaboura2025arbcomprehensivearabicmultimodal, title={ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark}, author={Sara Ghaboura and Ketan More and Wafa Alghallabi and Omkar Thawakar and Jorma Laaksonen and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer}, year={2025}, eprint={2505.17021}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2505.17021}, }
相关机构
- MBZUAI
- IVAL
- Oryx




