TAVGBench
收藏资源简介:
TAVGBench是由西北工业大学和上海人工智能实验室等机构合作开发的大规模数据集,包含超过170万个音频-视频对,总时长达到11.8万小时。该数据集来源于YouTube视频,通过自动化的粗到细文本标注流程进行详细描述,确保每个音频视频对都有详细的音频和视频内容描述。TAVGBench旨在支持文本到可听视频生成(TAVG)任务的研究,通过提供大规模、高质量的训练和测试数据,推动多模态生成技术的发展,特别是在需要同步音频和视频的场景中。
TAVGBench is a large-scale dataset co-developed by Northwestern Polytechnical University, Shanghai AI Laboratory and other institutions. It contains over 1.7 million audio-video pairs, with a total duration of 118,000 hours. Sourced from YouTube videos, each audio-video pair is equipped with detailed descriptions of both audio and visual content via an automated coarse-to-fine text annotation pipeline, ensuring the comprehensiveness of the content annotations. TAVGBench aims to support research on the Text-to-Audible Video Generation (TAVG) task. By providing large-scale, high-quality training and test datasets, it promotes the development of multimodal generation technologies, especially in scenarios requiring synchronized audio and video generation.




