重构的综合数据集
收藏资源简介:
本文重构了一个综合数据集,该数据集由北京大学多媒体信息处理国家重点实验室、北京科学院人工智能研究所、中国科学院大学人工智能学院和北京智源人工智能研究院共同创建,旨在评估视觉认知、几何理解和跨任务泛化能力。数据集覆盖了视觉计数、结构感知和空间变换三个核心领域,包含了高质量的步骤式推理数据,用于激活和增强视觉语言模型在视觉推理任务中的潜力。
This paper reconstructs a comprehensive dataset jointly created by the State Key Laboratory of Media Computing at Peking University, the Institute of Artificial Intelligence of Beijing Academy of Sciences, the School of Artificial Intelligence of the University of Chinese Academy of Sciences, and the Beijing Academy of Artificial Intelligence (BAAI). This dataset is designed to evaluate visual cognition, geometric understanding and cross-task generalization capabilities. It covers three core domains: visual counting, structural perception and spatial transformation, and contains high-quality step-by-step reasoning data to activate and enhance the potential of vision-language models in visual reasoning tasks.




