遇见数据集

tiiuae/visres_bench

收藏
Hugging Face2026-03-10 更新2026-05-10 收录
官方服务:

资源简介:

--- license: apache-2.0 task_categories: - visual-question-answering - image-to-text language: - en tags: - benchmark - vision - reasoning - multimodal - evaluation pretty_name: VisRes-Bench dataset_info: - config_name: level_1_global_occlusion_50 features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_global_occlusion_70 features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_global_occlusion_80 features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_edges features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_location_random_sampling features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_brightness features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_blur features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_rotation features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_rotation_random_sampling features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_edges_random_sampling features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_1_location features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 1000 - config_name: level_2_uniform_count features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_2_count_progression features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_2_uniform_orientation features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 458 - config_name: level_2_count_2_same_1_diff features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_2_orientation_2same_1diff features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 498 - config_name: level_2_uniform_color features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_2_count_arithmetic features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_2_count_minmax features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_2_orientation_3_diff features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_2_color_2same_1diff features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_2_color_3_diff features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_2_count_3_diff features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_3_spiral_color_orientation features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 350 - config_name: level_3_spiral_color_orientation features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 464 - config_name: level_3_coupled_color_count features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 500 - config_name: level_3_independent_color_object_orientation features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 355 - config_name: level_3_coupled_color_orientation features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 374 - config_name: level_3_Independent_count_object_color features: - name: id dtype: string - name: task dtype: string - name: level dtype: string - name: guided_question dtype: string - name: generic_question dtype: string - name: images sequence: image - name: question dtype: string - name: answer dtype: string splits: - name: test num_examples: 479 configs: - config_name: level_1_global_occlusion_50 data_files: - split: test path: level_1_global_occlusion_50percent/test-* - config_name: level_1_global_occlusion_70 data_files: - split: test path: level_1_global_occlusion_70percent/test-* - config_name: level_1_global_occlusion_80 data_files: - split: test path: level_1_global_occlusion_80percent/test-* - config_name: level_1_edges data_files: - split: test path: level_1_edges_eval_6k_location_only_dino_mode_options/test-* - config_name: level_1_location_random_sampling data_files: - split: test path: level_1_eval_6k_location_only_random_sampling/test-* - config_name: level_1_brightness data_files: - split: test path: level_1_eval_6k_brightness_dino_options/test-* - config_name: level_1_blur data_files: - split: test path: level_1_eval_6k_blur_dino_options/test-* - config_name: level_1_rotation data_files: - split: test path: level_1_eval_6k_rotation_direct_dino_options/test-* - config_name: level_1_rotation_random_sampling data_files: - split: test path: level_1_eval_6k_single_rotation_same_options/test-* - config_name: level_1_edges_random_sampling data_files: - split: test path: level_1_edges_eval_6k_location_only_random_sampling/test-* - config_name: level_1_location data_files: - split: test path: level_1_eval_6k_location_only_dino_mode_options/test-* - config_name: level_2_uniform_count data_files: - split: test path: level_2_count_only/test-* - config_name: level_2_count_progression data_files: - split: test path: level_2_count_progression_mixed/test-* - config_name: level_2_uniform_orientation data_files: - split: test path: level_2_orientation_only/test-* - config_name: level_2_count_2_same_1_diff data_files: - split: test path: level_2_count_distribution_2same_1diff/test-* - config_name: level_2_orientation_2same_1diff data_files: - split: test path: level_2_orientation_distribution_2same_1diff/test-* - config_name: level_2_uniform_color data_files: - split: test path: level_2_color_only/test-* - config_name: level_2_count_arithmetic data_files: - split: test path: level_2_count_operations/test-* - config_name: level_2_count_minmax data_files: - split: test path: level_2_count_minmax/test-* - config_name: level_2_orientation_3_diff data_files: - split: test path: level_2_orientation_distribution/test-* - config_name: level_2_color_2same_1diff data_files: - split: test path: level_2_color_distribution_2same_1diff/test-* - config_name: level_2_color_3_diff data_files: - split: test path: level_2_color_distribution/test-* - config_name: level_2_count_3_diff data_files: - split: test path: level_2_count_distribution/test-* - config_name: level_3_spiral_color_orientation data_files: - split: test path: level_3_compositional_spiral_orientation/test-* - config_name: level_3_spiral_color_orientation data_files: - split: test path: level_3_compositional_spiral_object_color/test-* - config_name: level_3_coupled_color_count data_files: - split: test path: level_3_coupled_count_color/test-* - config_name: level_3_independent_color_object_rientation data_files: - split: test path: level_3_independent_color_object_orientation/test-* - config_name: level_3_coupled_color_orientation data_files: - split: test path: level_3_coupled_orientation_color/test-* - config_name: level_3_Independent_count_object_color data_files: - split: test path: level_3_independent_distribution_arithmetic_object/test-* --- # VisRes Bench [![GitHub](https://img.shields.io/badge/GitHub-181717?style=flat-square&logo=github&logoColor=white)](https://visres-bench.github.io/) [![arXiv](https://img.shields.io/badge/arXiv-2512.21194-b31b1b?style=flat-square&logo=arxiv&logoColor=white)](https://arxiv.org/abs/2512.21194) **VisRes Bench** is a benchmark for evaluating the **visual reasoning** capabilities of Vision-Language Models (VLMs) in naturalistic settings without contextual language supervision. It is introduced in the paper [*VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs*](https://arxiv.org/abs/2512.21194). ## Paper Summary Vision-Language Models excel at captioning and VQA, but it is unclear how much they rely on **visual reasoning** versus **linguistic priors**. VisRes addresses this by using **image-only, four-choice** tasks on **real-world images** (~19,000 samples) so that performance reflects visual reasoning rather than textual shortcuts. The benchmark is organized in **three levels** of increasing complexity: - **Level 1 — Perceptual grounding:** Local patch completion (masked tile + 4 candidate patches) under perturbations (blur, brightness, rotation, edges, location) and global occlusion (50% or 80% of the image masked). Tests robustness and amodal completion. - **Level 2 — Single-attribute rule:** Raven-style 3×3 grids with one missing cell; one attribute (color, count, or orientation) follows a row-wise rule. Includes uniform, 3-different, 2-similar-1-different, count progression, arithmetic, and min-max subtasks (~5,956 samples). - **Level 3 — Multi-attribute composition:** Same 3×3 format but multiple attributes (color, count, orientation, object identity) with row-wise, grid-wise, or spiral rules (~2,522 samples). **Main findings:** State-of-the-art VLMs perform near **random (25%)** on many subtasks under subtle perceptual changes. Performance is stronger on color than count, and weakest on orientation. When the same logical structure is given as **text**, models do much better, indicating a **visual-to-symbolic** bottleneck rather than a pure reasoning limit. Higher resolution and guided/thinking prompts help but do not close the gap to human baselines. --- ## Main Results (Guided Prompting, Thinking Mode When Available) Accuracy (%) across levels and subtasks. Random chance = 25%. <table> <thead> <tr> <th>Setting</th> <th>GPT-5</th> <th>GPT-4o</th> <th>Gemini-2.5</th> <th>Qwen3-VL-4B</th> <th>Qwen3-VL-30B</th> <th>Mimo-VL-7B</th> </tr> </thead> <tbody> <tr><td colspan="7"><strong>Level-1</strong></td></tr> <tr><td>Edges</td><td>27.17</td><td>23.91</td><td>25.00</td><td>16.67</td><td>25.00</td><td>22.30</td></tr> <tr><td>Location</td><td>23.71</td><td>20.62</td><td>26.00</td><td>23.16</td><td>22.40</td><td>25.77</td></tr> <tr><td>Rotation</td><td>35.42</td><td>26.04</td><td>34.38</td><td>37.50</td><td>36.05</td><td>29.17</td></tr> <tr><td>Brightness</td><td>25.26</td><td>27.37</td><td>27.37</td><td>31.52</td><td>29.47</td><td>27.37</td></tr> <tr><td>Blur</td><td>31.18</td><td>25.26</td><td>26.32</td><td>24.73</td><td>24.28</td><td>26.32</td></tr> <tr><td>Global@50%</td><td>42.86</td><td>20.88</td><td>57.14</td><td>37.50</td><td>47.25</td><td>48.35</td></tr> <tr><td>Global@80%</td><td>32.61</td><td>22.83</td><td>36.96</td><td>25.88</td><td>35.87</td><td>30.43</td></tr> <tr><td><strong>Level-1 Average</strong></td><td><strong>31.10</strong></td><td><strong>23.86</strong></td><td><strong>33.28</strong></td><td><strong>28.17</strong></td><td><strong>31.20</strong></td><td><strong>29.22</strong></td></tr> <tr><td colspan="7"><strong>Level-2</strong></td></tr> <tr><td>Uniform Color</td><td>96.00</td><td>21.00</td><td>97.00</td><td>66.20</td><td>88.00</td><td>78.95</td></tr> <tr><td>Uniform Count</td><td>61.00</td><td>25.00</td><td>90.91</td><td>40.82</td><td>59.00</td><td>52.75</td></tr> <tr><td>Uniform Orientation</td><td>22.22</td><td>25.25</td><td>26.53</td><td>26.00</td><td>23.00</td><td>19.19</td></tr> <tr><td>Count Progression</td><td>50.00</td><td>13.00</td><td>77.00</td><td>37.20</td><td>48.00</td><td>36.96</td></tr> <tr><td>Count Arithmetic</td><td>52.00</td><td>22.00</td><td>75.76</td><td>43.20</td><td>49.00</td><td>33.33</td></tr> <tr><td><strong>Level-2 Average</strong></td><td><strong>49.79</strong></td><td><strong>24.12</strong></td><td><strong>62.29</strong></td><td><strong>37.18</strong></td><td><strong>46.75</strong></td><td><strong>39.15</strong></td></tr> <tr><td colspan="7"><strong>Level-3</strong></td></tr> <tr><td>Independent Color-Object-Orientation</td><td>34.00</td><td>25.25</td><td>38.00</td><td>27.39</td><td>32.60</td><td>19.00</td></tr> <tr><td>Independent Count-Object-Color</td><td>34.00</td><td>24.00</td><td>44.00</td><td>29.45</td><td>36.34</td><td>29.00</td></tr> <tr><td>Coupled Color-Orientation</td><td>24.24</td><td>24.00</td><td>16.33</td><td>26.13</td><td>29.43</td><td>20.00</td></tr> <tr><td>Coupled Color-Count</td><td>30.00</td><td>22.00</td><td>21.21</td><td>27.46</td><td>33.33</td><td>28.00</td></tr> <tr><td>Spiral Color-Count-Object</td><td>56.00</td><td>30.00</td><td>54.17</td><td>28.63</td><td>36.00</td><td>33.00</td></tr> <tr><td><strong>Level-3 Average</strong></td><td><strong>34.39</strong></td><td><strong>23.86</strong></td><td><strong>33.73</strong></td><td><strong>26.31</strong></td><td><strong>31.36</strong></td><td><strong>25.17</strong></td></tr> </tbody> </table> --- ## Finetuning on Level-1 (Qwen2.5-VL-3B) <table> <thead> <tr> <th>Setting</th> <th>Original</th> <th>Finetuned</th> <th>Human Baseline</th> </tr> </thead> <tbody> <tr><td>Location</td><td>24.3</td><td>42.8</td><td>94.1</td></tr> <tr><td>Blur</td><td>23.9</td><td>37.5</td><td>84.3</td></tr> <tr><td>Brightness</td><td>23.7</td><td>39.8</td><td>85.6</td></tr> <tr><td>Rotation</td><td>25.5</td><td>50.8</td><td>92.0</td></tr> <tr><td>Edges</td><td>25.1</td><td>33.2</td><td>82.6</td></tr> <tr><td>Global (50%)</td><td>24.9</td><td>52.2</td><td>96.1</td></tr> <tr><td>Global (80%)</td><td>23.9</td><td>38.6</td><td>98.0</td></tr> <tr><td><strong>Average</strong></td><td><strong>24.5</strong></td><td><strong>43.7</strong></td><td><strong>90.4</strong></td></tr> </tbody> </table> --- ## Single-Attribute Recognition (Perceptual Grounding) Accuracy (%) when models are asked to report a single attribute (color, orientation, or count) for one grid cell. <table> <thead> <tr> <th>Attribute</th> <th>GPT-4o</th> <th>GPT-5</th> </tr> </thead> <tbody> <tr><td>Color</td><td>84.6</td><td>97.6</td></tr> <tr><td>Orientation</td><td>39.8</td><td>49.6</td></tr> <tr><td>Count</td><td>72.4</td><td>94.2</td></tr> </tbody> </table> --- ## Impact of Thinking Mode Accuracy (%) with thinking mode enabled (✓) vs disabled (✗). Open-source models improve substantially with thinking. <table> <thead> <tr> <th>Level</th> <th>GPT-5 (high)</th> <th>GPT-5 (low)</th> <th>Mimo-VL ✓</th> <th>Mimo-VL ✗</th> <th>Qwen3-4B ✓</th> <th>Qwen3-4B ✗</th> <th>Qwen3-30B ✓</th> <th>Qwen3-30B ✗</th> </tr> </thead> <tbody> <tr><td>Level-1</td><td>32.61</td><td>31.43</td><td>29.22</td><td>23.91</td><td>28.17</td><td>23.16</td><td>31.20</td><td>23.60</td></tr> <tr><td>Level-2</td><td>49.79</td><td>47.01</td><td>39.15</td><td>26.68</td><td>37.18</td><td>24.08</td><td>46.75</td><td>28.25</td></tr> <tr><td>Level-3</td><td>34.39</td><td>32.89</td><td>25.17</td><td>25.23</td><td>26.31</td><td>23.50</td><td>31.36</td><td>24.00</td></tr> </tbody> </table> --- ## Impact of Image Resolution (GPT-5) Accuracy (%) at different input resolutions. All levels improve with higher resolution. <table> <thead> <tr> <th>Resolution</th> <th>Level-1</th> <th>Level-2</th> <th>Level-3</th> </tr> </thead> <tbody> <tr><td>512×512</td><td>45.17</td><td>42.83</td><td>31.63</td></tr> <tr><td>1024×1024</td><td>54.01</td><td>49.61</td><td>35.48</td></tr> <tr><td>2048×2048</td><td>56.51</td><td>48.99</td><td>40.07</td></tr> </tbody> </table> --- ## Citation ```bibtex @article{visres2025, title={VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs}, author={Malagurski T{\"o}rtei, Brigitta and Dahou, Yasser and Huynh, Ngoc Dung and others}, journal={arXiv preprint arXiv:2512.21194}, year={2025} } ```

许可证:Apache-2.0 任务类别: - 视觉问答(visual-question-answering) - 图像到文本(image-to-text) 语言: - 英语 标签: - 基准测试(benchmark) - 计算机视觉(vision) - 推理(reasoning) - 多模态(multimodal) - 评估(evaluation) 友好名称:VisRes-Bench 数据集信息: 以下为各配置的特征与划分结构(除特别说明外,其余配置特征与划分一致): - 配置名称:等级1全局遮挡50%(level_1_global_occlusion_50):测试集样本量1000 - 配置名称:等级1全局遮挡70%(level_1_global_occlusion_70):测试集样本量1000 - 配置名称:等级1全局遮挡80%(level_1_global_occlusion_80):测试集样本量1000 - 配置名称:等级1边缘任务(level_1_edges):测试集样本量1000 - 配置名称:等级1位置随机采样(level_1_location_random_sampling):测试集样本量1000 - 配置名称:等级1亮度扰动(level_1_brightness):测试集样本量1000 - 配置名称:等级1模糊扰动(level_1_blur):测试集样本量1000 - 配置名称:等级1旋转任务(level_1_rotation):测试集样本量1000 - 配置名称:等级1旋转随机采样(level_1_rotation_random_sampling):测试集样本量1000 - 配置名称:等级1边缘随机采样(level_1_edges_random_sampling):测试集样本量1000 - 配置名称:等级1位置任务(level_1_location):测试集样本量1000 - 配置名称:等级2均匀数量任务(level_2_uniform_count):测试集样本量500 - 配置名称:等级2数量递进任务(level_2_count_progression):测试集样本量500 - 配置名称:等级2均匀方向任务(level_2_uniform_orientation):测试集样本量458 - 配置名称:等级2两相似一不同数量任务(level_2_count_2_same_1_diff):测试集样本量500 - 配置名称:等级2两相似一不同方向任务(level_2_orientation_2same_1diff):测试集样本量498 - 配置名称:等级2均匀颜色任务(level_2_uniform_color):测试集样本量500 - 配置名称:等级2数量算术任务(level_2_count_arithmetic):测试集样本量500 - 配置名称:等级2数量极值任务(level_2_count_minmax):测试集样本量500 - 配置名称:等级2三种不同方向任务(level_2_orientation_3_diff):测试集样本量500 - 配置名称:等级2两相似一不同颜色任务(level_2_color_2same_1diff):测试集样本量500 - 配置名称:等级2三种不同颜色任务(level_2_color_3_diff):测试集样本量500 - 配置名称:等级2三种不同数量任务(level_2_count_3_diff):测试集样本量500 - 配置名称:等级3螺旋颜色-方向任务(level_3_spiral_color_orientation):测试集样本量350/464(两处配置) - 配置名称:等级3耦合颜色-数量任务(level_3_coupled_color_count):测试集样本量500 - 配置名称:等级3独立颜色-物体-方向任务(level_3_independent_color_object_orientation):测试集样本量355 - 配置名称:等级3耦合颜色-方向任务(level_3_coupled_color_orientation):测试集样本量374 - 配置名称:等级3独立数量-物体-颜色任务(level_3_Independent_count_object_color):测试集样本量479 配置文件: - 配置名称:等级1全局遮挡50%:数据文件路径为level_1_global_occlusion_50percent/test-*,划分:测试集 - 配置名称:等级1全局遮挡70%:数据文件路径为level_1_global_occlusion_70percent/test-*,划分:测试集 - 配置名称:等级1全局遮挡80%:数据文件路径为level_1_global_occlusion_80percent/test-*,划分:测试集 - 其余配置的数据文件路径与划分可参照原文路径格式,划分均为测试集 --- # VisRes基准测试(VisRes Bench) [![GitHub](https://img.shields.io/badge/GitHub-181717?style=flat-square&logo=github&logoColor=white)](https://visres-bench.github.io/) [![arXiv](https://img.shields.io/badge/arXiv-2512.21194-b31b1b?style=flat-square&logo=arxiv&logoColor=white)](https://arxiv.org/abs/2512.21194) **VisRes基准测试(VisRes Bench)** 是一款用于评估视觉语言模型(Vision-Language Models, VLMs)在无上下文语言监督的自然场景下视觉推理能力的基准测试集,相关研究论文为《VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs》。 ## 论文摘要 视觉语言模型在图像描述与视觉问答任务中表现优异,但目前尚不明确其性能在多大程度上依赖于视觉推理而非语言先验。VisRes基准测试通过在真实世界图像(约19000个样本)上采用仅基于图像的四选一任务,确保模型性能能够反映真实视觉推理能力而非文本捷径。 该基准测试按照复杂度递增的三个等级进行组织: - **等级1 — 感知基础任务**:在各类扰动(模糊、亮度变化、旋转、边缘处理、位置偏移)与全局遮挡(图像50%或80%区域被掩码)下的局部图像块补全任务(提供掩码图像块与4个候选图像块),用于测试模型的鲁棒性与无模态补全能力。 - **等级2 — 单属性规则任务**:带有一个缺失单元格的雷文式(Raven-style)3×3网格,单个属性(颜色、数量或方向)遵循行级规则。包含均匀分布、3种不同、2相似1不同、数量递进、算术运算与极值(min-max)等子任务,总计约5956个样本。 - **等级3 — 多属性组合任务**:同样采用3×3网格格式,但包含多个属性(颜色、数量、方向、物体标识),且遵循行级、网格级或螺旋规则,总计约2522个样本。 **核心发现**:当前主流的视觉语言模型在诸多带有细微感知变化的子任务上的准确率接近随机猜测水平(25%)。模型在颜色属性任务上的表现优于数量属性,而在方向属性任务上表现最差。当相同的逻辑结构以文本形式给出时,模型性能会大幅提升,这表明模型存在视觉到符号的转换瓶颈,而非纯粹的推理能力限制。更高的图像分辨率、引导式/思考型提示能够改善模型性能,但仍无法缩小与人类基准的差距。 --- ## 主实验结果(引导式提示,启用思考模式,若支持) 各等级与子任务的准确率(%),随机猜测准确率为25%。 | 设置 | GPT-5 | GPT-4o | Gemini-2.5 | Qwen3-VL-4B | Qwen3-VL-30B | Mimo-VL-7B | |---|---|---|---|---|---|---| | **等级1** | | | | | | | | 边缘任务 | 27.17 | 23.91 | 25.00 | 16.67 | 25.00 | 22.30 | | 位置任务 | 23.71 | 20.62 | 26.00 | 23.16 | 22.40 | 25.77 | | 旋转任务 | 35.42 | 26.04 | 34.38 | 37.50 | 36.05 | 29.17 | | 亮度任务 | 25.26 | 27.37 | 27.37 | 31.52 | 29.47 | 27.37 | | 模糊任务 | 31.18 | 25.26 | 26.32 | 24.73 | 24.28 | 26.32 | | 全局遮挡50% | 42.86 | 20.88 | 57.14 | 37.50 | 47.25 | 48.35 | | 全局遮挡80% | 32.61 | 22.83 | 36.96 | 25.88 | 35.87 | 30.43 | | **等级1平均准确率** | **31.10** | **23.86** | **33.28** | **28.17** | **31.20** | **29.22** | | **等级2** | | | | | | | | 均匀颜色任务 | 96.00 | 21.00 | 97.00 | 66.20 | 88.00 | 78.95 | | 均匀数量任务 | 61.00 | 25.00 | 90.91 | 40.82 | 59.00 | 52.75 | | 均匀方向任务 | 22.22 | 25.25 | 26.53 | 26.00 | 23.00 | 19.19 | | 数量递进任务 | 50.00 | 13.00 | 77.00 | 37.20 | 48.00 | 36.96 | | 数量算术任务 | 52.00 | 22.00 | 75.76 | 43.20 | 49.00 | 33.33 | | **等级2平均准确率** | **49.79** | **24.12** | **62.29** | **37.18** | **46.75** | **39.15** | | **等级3** | | | | | | | | 独立颜色-物体-方向任务 | 34.00 | 25.25 | 38.00 | 27.39 | 32.60 | 19.00 | | 独立数量-物体-颜色任务 | 34.00 | 24.00 | 44.00 | 29.45 | 36.34 | 29.00 | | 耦合颜色-方向任务 | 24.24 | 24.00 | 16.33 | 26.13 | 29.43 | 20.00 | | 耦合颜色-数量任务 | 30.00 | 22.00 | 21.21 | 27.46 | 33.33 | 28.00 | | 螺旋颜色-数量-物体任务 | 56.00 | 30.00 | 54.17 | 28.63 | 36.00 | 33.00 | | **等级3平均准确率** | **34.39** | **23.86** | **33.73** | **26.31** | **31.36** | **25.17** | --- ## 等级1微调实验(基于Qwen2.5-VL-3B) | 设置 | 原始模型 | 微调后模型 | 人类基准 | |---|---|---|---| | 位置任务 | 24.3 | 42.8 | 94.1 | | 模糊任务 | 23.9 | 37.5 | 84.3 | | 亮度任务 | 23.7 | 39.8 | 85.6 | | 旋转任务 | 25.5 | 50.8 | 92.0 | | 边缘任务 | 25.1 | 33.2 | 82.6 | | 全局遮挡50% | 24.9 | 52.2 | 96.1 | | 全局遮挡80% | 23.9 | 38.6 | 98.0 | | **平均准确率** | **24.5** | **43.7** | **90.4** | --- ## 单属性识别任务(感知基础) 当模型被要求报告单个网格单元格的单一属性(颜色、方向或数量)时的准确率(%)。 | 属性 | GPT-4o | GPT-5 | |---|---|---| | 颜色 | 84.6 | 97.6 | | 方向 | 39.8 | 49.6 | | 数量 | 72.4 | 94.2 | --- ## 思考模式的影响 启用思考模式(✓)与禁用思考模式(✗)下的准确率(%)。开源模型在启用思考模式后性能提升显著。 | 等级 | GPT-5(高算力) | GPT-5(低算力) | Mimo-VL ✓ | Mimo-VL ✗ | Qwen3-4B ✓ | Qwen3-4B ✗ | Qwen3-30B ✓ | Qwen3-30B ✗ | |---|---|---|---|---|---|---|---|---| | 等级1 | 32.61 | 31.43 | 29.22 | 23.91 | 28.17 | 23.16 | 31.20 | 23.60 | | 等级2 | 49.79 | 47.01 | 39.15 | 26.68 | 37.18 | 24.08 | 46.75 | 28.25 | | 等级3 | 34.39 | 32.89 | 25.17 | 25.23 | 26.31 | 23.50 | 31.36 | 24.00 | --- ## 图像分辨率的影响(基于GPT-5) 不同输入分辨率下的准确率(%)。所有等级的性能均随分辨率提升而改善。 | 分辨率 | 等级1 | 等级2 | 等级3 | |---|---|---|---| | 512×512 | 45.17 | 42.83 | 31.63 | | 1024×1024 | 54.01 | 49.61 | 35.48 | | 2048×2048 | 56.51 | 48.99 | 40.07 | --- ## 引用 bibtex @article{visres2025, title={VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs}, author={Malagurski Törtei, Brigitta and Dahou, Yasser and Huynh, Ngoc Dung and others}, journal={arXiv preprint arXiv:2512.21194}, year={2025} }

提供机构:
tiiuae
二维码
社区交流群
二维码
科研交流群
商业服务