Voxel51/VisualOverload
收藏资源简介:
VisualOverload是一个视觉问答(VQA)基准测试数据集,旨在评估视觉语言模型(VLMs)在密集、细节丰富的场景中的细粒度视觉理解能力。数据集包含2,720个问题-答案对,基于150幅公共领域绘画的高分辨率扫描图像(分辨率约为4K,如3840×2160)。这些图像被手动标注了问题,涵盖六个任务类别:活动识别、属性识别、计数、光学字符识别(OCR)、推理和场景理解。问题分为三种难度级别(简单、中等、困难)和三种回答类型(选择题、自由计数、自由OCR)。数据集假设当前基准测试可能高估了VLMs的性能,特别是在密集场景中,编码和推理细节仍具挑战性。在测试的37个模型中,最佳模型(o3)在最难子集上的准确率仅为19.6%,整体准确率为69.5%。错误分析揭示了包括弱计数、OCR失败和复杂任务下的逻辑不一致等失败模式。数据集仅用于评估,地面真实答案未公开,需通过官方评估服务器进行评分。语言为英语,许可证为CC BY-SA 4.0,图像为公共领域艺术品(CC0)。数据集由Paul Gavrikov等人策划,并由Voxel51转换为FiftyOne格式。
VisualOverload is a visual question answering (VQA) benchmark comprising 2,720 question–answer pairs with privately held ground-truth responses, designed to probe the fine-grained visual understanding of vision-language models (VLMs) in densely populated (overloaded) scenes. The dataset consists of 150 high-resolution scans of public-domain paintings, manually annotated with questions across six task categories: activity, attributes, counting, OCR, reasoning, and scene. Questions are categorized into three difficulty levels (easy, medium, hard) and three answer types (choice, freeform counting, freeform OCR). The authors hypothesize that current benchmarks overestimate VLM performance, and encoding and reasoning over details remain challenging in dense scenes. Among 37 tested models, the best model (o3) achieves only 19.6% accuracy on the hardest split and 69.5% overall. Error analysis reveals failure modes such as weak counting, OCR failures, and logical inconsistencies. The dataset is evaluation-only, with ground-truth answers withheld and accessible via an official evaluation server. The language is English, licensed under CC BY-SA 4.0, with images being royalty-free public-domain artwork (CC0). Curated by Paul Gavrikov et al. and shared in FiftyOne format by Voxel51.




