LongViTU
收藏资源简介:
LongViTU是一个用于长视频理解的大规模数据集,由北京大学、BIGAI和新加坡国立大学的研究团队共同创建。该数据集包含约121k个高质量的问答对,覆盖约900小时的视频内容,平均每个视频的问答对时长为4.6分钟。数据集通过自动生成的层次化树结构构建,确保了问答对的高质量和时间戳的精确标注。数据集的内容涵盖了多样化的真实世界场景,适用于长视频和流媒体视频的理解任务,旨在解决现有数据集在时间标注、场景多样性和问答精确性方面的不足。LongViTU的应用领域包括视频问答、长视频理解以及流媒体视频分析等。
LongViTU is a large-scale dataset for long-form video understanding, jointly created by research teams from Peking University, BIGAI, and the National University of Singapore. This dataset contains approximately 121,000 high-quality question-answer pairs, covering around 900 hours of video content, with the average duration of the question-answer pairs per video being 4.6 minutes. The dataset is constructed using automatically generated hierarchical tree structures, which ensures the high quality of the question-answer pairs and the precision of timestamp annotations. Its content covers diverse real-world scenarios, and is applicable to long-form video and streaming video understanding tasks, aiming to address the limitations of existing datasets in terms of temporal annotation, scene diversity, and the accuracy of question-answer pairs. The application fields of LongViTU include video question answering, long-form video understanding, streaming video analysis, and other related fields.




