Event-Image/Video Pairs
收藏资源简介:
Event-Image/Video Pairs数据集由香港科技大学(广州)、鲁汶大学和香港中文大学联合创建,包含约140万条高质量的事件-图像/视频-文本配对数据。该数据集旨在通过多模态大语言模型(MLLM)实现对事件流的深度语义理解。数据集涵盖了多种场景,如驾驶场景和人体运动场景,数据来源包括静态图像、动态场景和人体运动视频。数据集的创建过程利用了开源的MLLM模型进行标注,并通过人工检查确保数据质量。该数据集主要用于事件描述生成、场景理解等任务,旨在解决事件数据在细粒度语义理解上的瓶颈问题。
The Event-Image/Video Pairs dataset was jointly created by The Hong Kong University of Science and Technology (Guangzhou), KU Leuven, and The Chinese University of Hong Kong. It contains approximately 1.4 million high-quality event-image/video-text paired samples. The dataset is designed to enable deep semantic understanding of event streams via multimodal large language models (MLLMs). It covers diverse scenarios including driving scenarios and human motion scenarios, with data sources including static images, dynamic scenes, and human motion videos. During its construction, open-source MLLM models were employed for annotation, and manual inspections were conducted to ensure data quality. This dataset is primarily used for tasks such as event caption generation and scene understanding, with the goal of addressing the bottleneck issue in fine-grained semantic understanding of event data.



