CelebV-Text
收藏资源简介:
CelebV-Text是由新加坡南洋理工大学S-Lab创建的大规模面部文本-视频数据集,包含70,000个野生面部视频片段,每个视频片段配以20个文本描述。数据集旨在促进面部文本到视频生成任务的研究,涵盖了从静态到动态的多种面部属性描述,如外观、动作和情感。创建过程包括数据收集、处理、标注和半自动文本生成,确保了文本与视频之间的高度相关性。该数据集适用于开发和评估面部文本到视频生成的先进模型,解决视频生成中缺乏高质量和相关文本描述的问题。
CelebV-Text is a large-scale facial text-video dataset developed by S-Lab at Nanyang Technological University, Singapore. It contains 70,000 unconstrained facial video clips, with each clip paired with 20 textual descriptions. This dataset aims to advance research on facial text-to-video generation tasks, covering diverse facial attribute descriptions ranging from static to dynamic types such as appearance, actions and emotions. Its construction workflow includes data collection, processing, annotation and semi-automatic text generation, which ensures a high degree of alignment between the textual descriptions and corresponding videos. This dataset is applicable to developing and evaluating state-of-the-art models for facial text-to-video generation, addressing the problem of insufficient high-quality and relevant textual descriptions in video generation research.




