STORM
收藏资源简介:
STORM数据集是一个全面的视觉评分数据集,旨在促进多模态大型语言模型(MLLMs)的可信排序回归能力。该数据集涵盖了五个常见视觉评分领域中的14个排序回归数据集,包含655K图像级配对及其对应的精心策划的视觉问答(VQA)。STORM数据集通过一个粗到精的处理流程,动态地考虑标签候选并提供可解释的思考,为MLLMs提供了一个通用且可信的排序思维范例。该数据集的创建过程包括收集和整合来自不同领域的多个数据集,并对数据进行标注和转换,使其适用于视觉评分任务。STORM数据集的应用领域包括图像质量评估、面部年龄估计、医学图像分级等,旨在解决MLLMs在视觉评分能力方面的不足,并为其提供一个全面的评估框架。
The STORM dataset is a comprehensive visual scoring dataset designed to enhance the trustworthy ranking regression capabilities of multimodal large language models (MLLMs). It covers 14 ranking regression datasets across five common visual scoring domains, containing 655K image-level pairs and their corresponding carefully curated visual question answering (VQA) content. The STORM dataset provides a generalizable and trustworthy ranking reasoning paradigm for MLLMs through a coarse-to-fine processing pipeline that dynamically considers label candidates and offers explainable reasoning. The development of the STORM dataset involves collecting and integrating multiple datasets from diverse domains, as well as annotating and transforming the data to make it suitable for visual scoring tasks. Application scenarios of the STORM dataset include image quality assessment, facial age estimation, medical image grading, and others, aiming to address the shortcomings of MLLMs in visual scoring capabilities and provide a comprehensive evaluation framework for such models.
STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset
作者信息
- Jinhong Wang1*, Shuo Tong1*, Jian Liu2, Dongqi Tang2, Haochao Ying1, Hongxia Xu1, Danny Chen3, Jintai Chen4, Jian Wu1✝
- 机构:1ZJU, 2Ant Group, 3University of Notre Dame, 4HKUST (Guangzhou)
- 状态:Underreview of NIPS 2025
- 通讯作者:Jian Wu
数据集概述
- 名称:STORM
- 目标:评估多模态大语言模型(MLLMs)在视觉评级任务中的表现
- 特点:
- 包含14个有序回归数据集,覆盖5个常见视觉评级领域
- 655K图像-标签对及精心设计的VQAs
- 提供粗到细的处理流程,动态考虑标签候选并提供可解释的思路
关键组件
- 广泛领域数据(14个数据集,5个领域)
- 多样化级别标注
- 粗到细的CoT(Chain-of-Thought)
- 一体化视觉评级框架
数据集用途
- 评估MLLMs在需要理解评级标签基本有序关系的场景中的表现
- 支持进一步研究MLLMs在视觉评级任务中的性能优化
资源链接
引用信息
bibtex @misc{wang2025stormbenchmarkingvisualrating, title={STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset}, author={Jinhong Wang and Shuo Tong and Jian liu and Dongqi Tang and Jintai Chen and Haochao Ying and Hongxia Xu and Danny Chen and Jian Wu}, year={2025}, eprint={2506.01738}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2506.01738}, }




