QCRI/BloomBench
收藏资源简介:
Almieyar-Oryx-BloomBench是Almieyar基准系列的一部分,是首个基于认知人类基础的双语(英语-阿拉伯语)多模态视觉语言模型基准。该基准基于Blooms Taxonomy,通过精心设计的图像-问题-答案任务,系统评估六个层次的认知能力:从基本的感知回忆到高阶的创造性合成。它旨在解决现有基准在评估认知弱点和提供针对性改进见解方面的不足。关键发现表明,当前先进的视觉语言模型存在认知不对称性,在语义理解方面表现强劲,但在事实回忆和创造性合成方面较弱;且阿拉伯语性能持续落后于英语,即使在强大的多语言模型中也是如此。数据集包含7,747个问答对,涵盖106个不同的叶子类别,采用4项选择题格式,图像来源于网络爬取的真实世界图像,质量验证率达到98.45%。
BloomBench is part of the Almieyar benchmarking series — the first cognitively human-grounded, bilingual (English–Arabic) multimodal benchmark for Vision-Language Models (VLMs). Grounded in Blooms Taxonomy, it systematically evaluates six levels of cognition through carefully designed image–question–answer tasks. It addresses the gap in existing benchmarks that obscure cognitive weaknesses and provide little insight for targeted improvement. A key finding reveals that state-of-the-art VLMs show a sharp cognitive asymmetry — strong in semantic understanding but substantially weaker in factual recall and creative synthesis, with Arabic performance consistently lagging English even in strong multilingual models. The dataset includes 7,747 QA pairs across 106 distinct leaf categories, in a 4-choice MCQ format, with images sourced from web-crawled real-world images and a quality validation rate of 98.45%.




