MultiTalk
收藏资源简介:
MultiTalk数据集是由韩国科学技术院创建,包含超过420小时的2D视频,涵盖20种不同语言,旨在解决多语言环境下3D说话头生成的问题。该数据集通过自动化管道从YouTube收集,每段视频都配有语言标签和伪转录,部分视频还包含伪3D网格顶点。数据集的创建过程包括视频收集、主动说话者验证和正面人脸验证,确保数据质量。MultiTalk数据集的应用领域主要集中在提升多语言3D说话头生成的准确性和表现力,通过引入语言特定风格嵌入,使模型能够捕捉每种语言独特的嘴部运动。
The MultiTalk dataset was developed by the Korea Advanced Institute of Science and Technology (KAIST). It contains over 420 hours of 2D videos spanning 20 distinct languages, and is designed to address the challenges of 3D talking head generation in multilingual scenarios. Collected from YouTube via an automated pipeline, each video in the dataset is paired with language labels and pseudo-transcripts, and a portion of the videos additionally include pseudo-3D mesh vertices. The dataset creation process encompasses video collection, active speaker verification, and frontal face verification to guarantee data quality. The primary application scenarios of the MultiTalk dataset focus on improving the accuracy and expressiveness of multilingual 3D talking head generation: by introducing language-specific style embeddings, models can capture the unique lip movements of each language.

- 1MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset韩国科学技术院 · 2024年



