TopicVD
收藏资源简介:
TopicVD是一个基于主题的视频支持的多模态机器翻译数据集,旨在推动纪录片翻译领域的研究。数据集由256部纪录片组成,总时长285小时,包含122,930对中英平行字幕,分为8个主题:经济、食物、历史、人物、军事、自然、社会和技术。数据集保留了每个视频字幕对的上下文信息,以支持利用纪录片的全局上下文进行视频引导的多模态机器翻译研究。数据集的建设过程包括数据收集、字幕处理、视频处理和数据构建等步骤。TopicVD旨在解决现有多模态机器翻译数据集在视频数据广泛性和主题多样性方面的不足,为纪录片翻译提供更丰富的视觉和文本信息。数据集的创建和应用对于推动视频引导的多模态机器翻译研究具有重要意义。
TopicVD is a topic-based video-supported multimodal machine translation dataset aimed at advancing research in the field of documentary translation. The dataset consists of 256 documentaries with a total duration of 285 hours, containing 122,930 pairs of Chinese-English parallel subtitles, and is categorized into 8 themes: economy, food, history, people, military, nature, society, and technology. The dataset retains the contextual information of each video-subtitle pair, to support research on video-guided multimodal machine translation that leverages the global contextual information of documentaries. The construction process of TopicVD includes steps such as data collection, subtitle processing, video processing and dataset building. TopicVD aims to address the shortcomings of existing multimodal machine translation datasets in terms of the breadth of video data and thematic diversity, providing richer visual and textual information for documentary translation. The creation and application of this dataset are of great significance for advancing research on video-guided multimodal machine translation.
TopicVD数据集概述
基本信息
- 数据集名称:TopicVD
- 托管平台:GitHub
- 托管地址:https://github.com/JinzeLv/TopicVD
数据集描述
(根据提供的README文件内容,该数据集未包含具体描述信息)

- 1TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries深圳大学应用技术学院, 深圳科技大学大数据与互联网学院 · 2025年



