VLM@school
收藏资源简介:
VLM@school 是一个用于评估视觉语言模型(VLMs)在德语环境中将视觉推理与学科特定背景知识相结合的能力的基准数据集。与广泛使用的英语基准相比,该数据集从九个领域(包括数学、历史、生物学和宗教)的真实中学课程中获取数据。基准包括超过 2,000 个开放式问题,这些问题的依据是 486 张图像,确保模型必须将视觉解释与事实推理相结合,而不是依赖于表面的文本线索。该数据集和评估协议作为严格的测试平台,旨在更好地理解和提高未来 AI 系统的视觉和语言推理能力。
VLM@school is a benchmark dataset designed to evaluate the ability of vision-language models (VLMs) to combine visual reasoning with domain-specific background knowledge in a German-language context. Compared to widely used English benchmarks, this dataset draws data from real secondary school curricula across nine domains including mathematics, history, biology and religion. The benchmark comprises over 2,000 open-ended questions grounded in 486 images, ensuring that models must integrate visual interpretation with factual reasoning rather than relying on superficial textual cues. This dataset and its evaluation protocol serve as a rigorous testbed aimed at better understanding and enhancing the visual and language reasoning capabilities of future AI systems.
- 1VLM@school -- Evaluation of AI image understanding on German middle school knowledge霍夫应用科学大学 · 2025年



