电视节目片段检索
收藏资源简介:
电视节目片段检索(TVR)是一个多模态数据集,最初用于视频语料库时刻检索,视频与自动语音识别生成的字幕配对。它包含从6个电视节目中收集的21.8K个视频,每个视频都与描述视频中特定片段的5个自然语言句子相关联。本数据集使用17,435个视频(包含87,175个片段)进行训练,并使用2,179个视频(包含10,895个片段)进行测试。
Television Program Segment Retrieval (TVR) is a multimodal dataset initially proposed for video corpus moment retrieval, which pairs videos with subtitles generated via automatic speech recognition (ASR). It contains 21.8K videos collected from 6 television programs, with each video associated with 5 natural language sentences that describe specific segments within the corresponding video. The dataset is split into training and test subsets: the training subset includes 17,435 videos (totaling 87,175 segments), while the test subset consists of 2,179 videos (totaling 10,895 segments).




