HuMo
收藏资源简介:
HuMo数据集是一个高质量的多模态数据集,包含文本、参考图像和音频三种模态的配对三元组条件。数据集构建经过两阶段的多模态数据处理流程,首先从大规模视频样本中检索与视频语义相同但视觉属性不同的参考图像,其次通过语音增强和语音-唇对齐估计进一步筛选具有同步音频轨道的视频样本。数据集的构建为后续多模态视频生成模型的学习提供了坚实的基础,旨在解决人本视频生成中数据稀缺和多模态协同控制困难的问题。
The HuMo dataset is a high-quality multimodal dataset that contains paired triplet conditions across three modalities: text, reference images, and audio. Its construction follows a two-stage multimodal data processing pipeline: first, retrieve reference images that share identical semantic content with the source video but possess distinct visual attributes from a large-scale video corpus; second, further filter video samples with synchronized audio tracks through speech enhancement and speech-lip alignment estimation. The dataset lays a solid foundation for the training of downstream multimodal video generation models, aiming to tackle the challenges of data scarcity and difficulties in multimodal collaborative control in human-centric video generation.

- 1通过清华大学,字节跳动智能创作实验室 · 2025年



