HRScene
收藏资源简介:
HRScene是一个用于高分辨率图像(HRI)理解的新颖统一基准,包含丰富的场景。它包含了25个真实世界的数据集和2个合成诊断数据集,分辨率范围从1,024 × 1,024到35,503 × 26,627。HRScene由10名研究生级别的标注员收集和重新标注,涵盖了从显微镜到放射学图像、街景、长距离照片和望远镜图像的25个场景。它包括真实世界对象的高分辨率图像、扫描文档和复合多图像。两个诊断评估数据集是通过将目标图像与金标准答案和干扰图像以不同顺序组合来合成的,以评估模型如何利用HRI中的区域。我们进行了广泛的实验,涉及28个VLMs,包括Gemini 2.0 Flash和GPT-4o。HRScene上的实验表明,当前的VLMs在现实世界任务上平均准确率约为50%,揭示了HRI理解中的重大差距。合成数据集上的结果表明,VLMs难以有效利用HRI区域,显示出显著的区域分化和迷失在中部的问题,为未来的研究提供了启示。
HRScene is a novel unified benchmark for high-resolution image (HRI) understanding, encompassing diverse scenarios. It consists of 25 real-world datasets and 2 synthetic diagnostic datasets, with resolutions ranging from 1,024 × 1,024 to 35,503 × 26,627. HRScene was collected and re-annotated by 10 graduate-student-level annotators, covering 25 scenario types spanning microscopic images, radiological images, street views, long-distance photographs, and telescopic images. Its corpus includes high-resolution images of real-world objects, scanned documents, and composite multi-image datasets. The two synthetic diagnostic datasets are constructed by combining target images, gold standard annotations, and distractor images in varying orders, aiming to evaluate how models leverage regional information within HRIs. We conducted extensive experiments involving 28 vision-language models (VLMs), including Gemini 2.0 Flash and GPT-4o. Experimental results on HRScene demonstrate that current VLMs achieve an average accuracy of approximately 50% on real-world tasks, revealing significant gaps in HRI understanding. Results on the synthetic datasets show that VLMs struggle to effectively utilize regional information in HRIs, exhibiting notable regional discrimination issues and the problem of getting trapped in central regions, which provides valuable insights for future research.
HRScene数据集概述
数据集简介
- 名称: HRScene
- 目的: 评估视觉大语言模型(VLMs)在高分辨率图像(HRI)理解方面的能力
- 特点:
- 首个统一的HRI理解基准
- 包含丰富场景的高分辨率图像
- 分辨率范围: 1,024×1,024至35,503×26,627像素
数据集构成
- 来源:
- 25个真实世界数据集
- 2个合成诊断数据集
- 标注:
- 由10名研究生级别标注员重新标注
- 覆盖25种场景(从显微镜到望远镜图像)
- 样本量: 7,081张图像(其中2,008张重新标注)
数据集分类
- 8个任务类别:
- Daily pictures
- Urban planning
- Paper scanned images
- Artwork
- Multi-subimages
- Remote sensing
- Medical Diagnosing
- Research understanding
数据集划分
- val: 750个样本(与人工标注样本相同)
- testmini: 1,000个样本(来自各真实世界数据集)
- test: 5,323个样本(答案标签不公开)
评估结果
- 模型表现:
- 当前VLMs平均准确率约50%
- 在合成数据集上表现差距超过20%
- 最佳表现模型: Qwen2-VL 72B(62.03%准确率)
诊断数据集发现
- 区域发散现象: 模型在不同图像区域表现不一致
- 曼哈顿距离现象: 性能随距离呈U型变化
数据访问
- 下载地址: Hugging Face Dataset
- 评估平台: 提供在线评估服务
模型提交
- 提交方式: 通过EvalAI平台
- 工具支持: 提供自动模型预测和提交管道
关键统计
- 图像分辨率分布: 从1k到超过35k像素
- 任务类型分布: 覆盖多种视觉上下文类型
- 问题长度分布: 记录每个问题的单词数分布




