CV-CapBench
收藏资源简介:
CV-CapBench是一个全面视觉字幕基准,由阿里巴巴集团和上海交通大学共同创建,旨在评估多模态大型语言模型在视觉字幕任务上的性能。该数据集包含6个视角和13个维度,涵盖了静态和动态视觉元素,旨在评估字幕的准确性和覆盖范围。数据集通过预标注和人工校正的方式构建,包含近1000张图片/视频每个维度,用于训练和评估多模态大型语言模型在视觉字幕方面的性能。该数据集的应用领域是计算机视觉,旨在解决视觉字幕任务中多模态大型语言模型的性能评估问题。
CV-CapBench is a comprehensive visual captioning benchmark co-created by Alibaba Group and Shanghai Jiao Tong University, aiming to evaluate the performance of multimodal large language models (LLMs) on visual captioning tasks. This dataset covers 6 perspectives and 13 dimensions, encompassing both static and dynamic visual elements, and is designed to assess the accuracy and coverage of generated captions. Constructed through pre-annotation and manual correction, the dataset contains nearly 1,000 images/videos per dimension, which is used for training and evaluating multimodal LLMs in visual captioning tasks. Its application domain is computer vision, and it aims to address the performance evaluation issue of multimodal LLMs in visual captioning tasks.




