SOU
收藏资源简介:
## 🌟 Highlights * **First-Ever Benchmark**: To the best of our knowledge, MLLMs based Small Object Understanding (SOU) tasks are proposed for the first time. A comprehensive benchmark (**SOUBench**), including relative datasets and baselines, is reported for the specific task. SOUBench fully reveals the shortcomings of current MLLMs in understanding small objects. * **Comprehensive Evaluation**: We design an effective automatic visual question-answer generation pipeline and introduce a comprehensive SOU-VQA evaluation dataset for small object understanding tasks, with **18,204** pairs and six relevant sub-tasks. Comprehensive experiments and comparisons are conducted in 15 state-of-the-art MLLMs to evaluate the small object understanding capability of MLLMs. Sufficient results reveal that current MLLMs have a weak understanding ability in the proposed tasks, even the best MLLM is still behind Human performance by 23.53%. * **Effcitive Fine-tuning**: We further construct **SOU-Train**, a multimodal VQA training dataset with **11,226** fine-grained annotations, to supervise the fine-tuning of the latest MLLM. The result denotes that the SOU-Train can effectively improve the small understanding ability of MLLM in different scenarios. Our research provides a crucial empirical foundation for the enhancement of the small object understanding capabilities of MLLMs. --- ## 📝 Citation If you find this project useful, please consider citing our work: ```bash @article{han2026can, title={Can Multimodal Large Language Models Truly Understand Small Objects?}, author={Han, Fujun and Chen, Junan and Zhu, Xintong and Ye, Jingqi and Mao, Xuanjie and Chen, Tao and Ye, Peng}, journal={arXiv preprint arXiv:2604.22884}, year={2026} } ``` license: Apache License 2.0 --- 数据集文件元信息以及数据文件,请浏览“数据集文件”页面获取。 当前数据集卡片使用的是默认模版,数据集的贡献者未提供更加详细的数据集介绍,但是您可以通过如下GIT Clone命令,或者ModelScope SDK来下载数据集 #### 下载方法 :modelscope-code[]{type="sdk"} :modelscope-code[]{type="git"}



