TVR
收藏资源简介:
TVR是一个大规模的多模态时刻检索数据集,由北卡罗来纳大学教堂山分校创建。该数据集包含109,000个查询,涉及21,800个来自6个不同类型电视节目的视频,每个查询与一个紧密的时间窗口相关联。TVR要求系统理解视频及其关联的字幕文本,使其更贴近现实。数据集还标注了查询类型,指示每个查询与视频、字幕或两者的关联程度,以便进行深入分析。通过严格的资格和后标注验证测试,确保了数据质量。此外,TVR还扩展了时刻检索任务,使其在多模态设置中更加现实,需要同时考虑视频和字幕文本。
TVR is a large-scale multimodal moment retrieval dataset created by the University of North Carolina at Chapel Hill. It contains 109,000 queries associated with 21,800 videos sourced from 6 genres of television programs, where each query is linked to a tight temporal window. TVR mandates that systems comprehend both the video content and its accompanying subtitle text, rendering the task more grounded in real-world scenarios. The dataset also annotates query categories, indicating the degree to which each query relates to the video, subtitles, or both, to enable in-depth analysis. Data quality is ensured through rigorous qualification and post-annotation validation tests. Furthermore, TVR expands the moment retrieval task to be more realistic in multimodal settings, requiring simultaneous consideration of both video and subtitle text.
- 1TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval北卡罗来纳大学教堂山分校 · 2020年



