Video Scene Graph Reasoning (VSGR)
收藏资源简介:
Video Scene Graph Reasoning (VSGR) 数据集由阿肯色大学和俄亥俄州立大学联合创建,旨在解决视频场景理解中的复杂关系和推理问题。该数据集包含190万帧视频,涵盖第三人称、第一人称和无人机视角,支持场景图生成、场景图预测、视频问答、视频字幕生成和关系推理五项任务。数据集的创建过程结合了实体场景图和过程图,通过超图结构捕捉多对象间的空间和时间关系。VSGR数据集的应用领域广泛,包括自动驾驶、智能监控、人机交互和多媒体分析等,旨在提升多模态大语言模型在动态视频场景中的理解和推理能力。
The Video Scene Graph Reasoning (VSGR) dataset was co-created by the University of Arkansas and The Ohio State University, aiming to address complex relational and reasoning challenges in video scene understanding. This dataset contains 1.9 million video frames, covering third-person, first-person, and drone perspectives, and supports five tasks including scene graph generation, scene graph prediction, video question answering, video caption generation, and relational reasoning. The construction of the VSGR dataset integrates entity scene graphs and procedural graphs, capturing spatial and temporal relationships among multiple objects via hypergraph structures. The VSGR dataset has broad application prospects in fields such as autonomous driving, intelligent surveillance, human-computer interaction and multimedia analysis, and is designed to enhance the understanding and reasoning capabilities of multimodal large language models in dynamic video scenes.




