ID-LoRA-TalkVid
收藏资源简介:
TalkVid Preprocessed for ID-LoRA 是一个专为身份驱动的音视频个性化任务设计的数据集,适用于文本到视频和文本到音频任务。数据集基于 TalkVid 源数据集,包含 5,796 个训练对、5,803 个独特视频片段和 600 位不同说话者。视频分辨率为 512x512,帧率为 25 fps,每个剪辑包含 121 帧(约 4.84 秒)。数据集提供了预计算的视频和音频潜在表示(VAE latents),以及结构化标注(包含视觉、语音、声音和文本四个部分)。每个训练对包括目标视频片段和参考视频片段(来自同一说话者),用于模型学习在保持说话者身份的同时生成音视频内容。数据集还包含说话者身份聚类信息和文本嵌入(Gemma 3 生成)。适用于训练 ID-LoRA 适配器,支持音视频联合生成任务。
TalkVid Preprocessed for ID-LoRA is a dataset specifically designed for identity-driven audio-visual personalization tasks, applicable to text-to-video and text-to-audio generation tasks. Based on the original TalkVid dataset, it contains 5,796 training pairs, 5,803 unique video clips and 600 distinct speakers. Each video clip has a resolution of 512×512, a frame rate of 25 fps, and consists of 121 frames (approximately 4.84 seconds). The dataset provides pre-computed video and audio latent representations (VAE latents), as well as structured annotations covering four aspects: visual, speech, audio and text. Each training pair includes a target video clip and a reference video clip from the same speaker, enabling models to learn to generate audio-visual content while retaining the speaker's identity. Additionally, the dataset contains speaker identity clustering information and text embeddings generated by Gemma 3. It is suitable for training ID-LoRA adapters and supports joint audio-visual generation tasks.



