TVQA
收藏资源简介:
TVQA是由希伯来大学计算机科学学院创建的大型视频问答数据集,包含150,000个问答对,覆盖6,500个视频片段,涉及多种主题如物体识别、场景理解和故事理解。数据集设计包括视频帧、字幕和语音三种模态信息,旨在通过多模态信息解决复杂问题。创建过程中,研究者采用人工标注和分类工具相结合的方法,分析各模态的重要性及其在数据集中的表现。TVQA的应用领域主要集中在评估和提升多模态AI模型的性能,特别是在需要综合视觉、听觉和文本信息的场景中。
TVQA is a large-scale video question answering (QA) dataset developed by the School of Computer Science at The Hebrew University of Jerusalem. It encompasses 150,000 question-answer pairs, spanning 6,500 video clips, and covers diverse topics including object recognition, scene understanding, and story comprehension. The dataset incorporates three modalities: video frames, subtitles, and speech, with the goal of addressing complex questions via multimodal information. During its development, researchers adopted a hybrid method combining manual annotation and classification tools, and analyzed the importance of each modality and its performance within the dataset. The main application fields of TVQA focus on evaluating and improving the performance of multimodal AI models, especially in scenarios that require comprehensive integration of visual, auditory and textual information.




