BigVideo
收藏资源简介:
该数据集名为BigVideo,是一个大规模的视频字幕翻译数据集,包含了450万句对以及总计9,981小时的视频资料。这些数据来源于YouTube和西瓜视频,旨在推动多模态机器翻译的发展。该数据集不仅包含了英汉两种语言的人工撰写字幕,而且重点筛选了高质量的视频-字幕配对。此外,数据集还包含了两个测试集:模糊测试集和明确测试集,它们被设计用来评估在翻译过程中视觉上下文的必要性。该数据集的规模达到了450万句对和9,981小时的视频资料,其任务专注于视频字幕翻译。
The dataset named BigVideo is a large-scale video subtitle translation dataset containing 4.5 million sentence pairs and a total of 9,981 hours of video footage. The data are sourced from YouTube and Xigua Video, with the goal of advancing research in multimodal machine translation. This dataset not only includes manually authored subtitles in both English and Chinese, but also places emphasis on curating high-quality video-subtitle pairs. Furthermore, the dataset comprises two test sets: the ambiguous test set and the explicit test set, which are developed to evaluate the necessity of incorporating visual context during the translation process. With a scale of 4.5 million sentence pairs and 9,981 hours of video content, this dataset is dedicated to the video subtitle translation task.




