ST-Align
收藏资源简介:
ST-Align数据集由北京航空航天大学、合肥工业大学、中国科学院信息工程研究所和美团联合创建,旨在支持细粒度的时空多模态理解任务。该数据集包含430万条训练样本,涵盖了15种不同的任务类型,数据来源包括WebVid-10M、Panda-70M、InternVid-10M等多个公开数据集以及自收集数据。数据集的创建过程包括内容对齐、坐标对齐和多任务指令调优三个阶段,确保模型能够逐步学习时空对齐和多任务处理能力。ST-Align数据集的应用领域包括时空视频定位、事件定位与描述、空间视频定位等,旨在解决现有多模态大语言模型在时空细粒度理解任务中的不足。
The ST-Align dataset, jointly created by Beijing University of Aeronautics and Astronautics, Hefei University of Technology, Institute of Information Engineering, Chinese Academy of Sciences, and Meituan, is designed to support fine-grained spatiotemporal multimodal understanding tasks. The dataset contains 4.3 million training samples, covering 15 different task types, and incorporates data from multiple publicly available datasets such as WebVid-10M, Panda-70M, InternVid-10M, as well as self-collected data. The dataset creation process involves three stages: content alignment, coordinate alignment, and multi-task instruction optimization, to ensure that the model can progressively learn spatiotemporal alignment and multi-task processing capabilities. The application fields of the ST-Align dataset include spatiotemporal video localization, event localization and description, and spatial video localization, aiming to address the limitations of existing multimodal large language models in understanding tasks at the spatiotemporal fine granularity.




