Allo-AVA
收藏资源简介:
Allo-AVA是由乔治亚理工学院创建的一个大规模多模态对话AI数据集,专门用于以第三人称视角的虚拟环境中的文本和音频驱动的虚拟形象手势动画。该数据集包含约1250小时的多样化视频内容,涵盖音频、转录文本和提取的关键点。数据集通过精确的时间戳映射,实现了语音与身体和面部手势的同步。Allo-AVA的多样性体现在其广泛的演讲者人口统计数据上,涵盖了多种年龄、性别和种族背景。该数据集旨在解决现有数据集在语音、面部表情和身体动作同步方面的不足,推动从虚拟现实到数字助手等应用领域的自然和上下文感知的虚拟形象动画模型的开发。
Allo-AVA is a large-scale multimodal conversational AI dataset created by the Georgia Institute of Technology, specifically tailored for text and audio-driven virtual avatar gesture animation in third-person virtual environments. This dataset encompasses roughly 1,250 hours of diverse video content, including audio, transcribed text, and extracted keypoints. Leveraging precise timestamp mapping, it achieves seamless synchronization between speech and both bodily and facial gestures. The diversity of Allo-AVA is reflected in its broad speaker demographics, which cover a wide range of age, gender, and racial backgrounds. This dataset aims to address the shortcomings of existing datasets regarding the synchronization of speech, facial expressions, and bodily movements, and advance the development of natural and context-aware virtual avatar animation models for applications spanning from virtual reality to digital assistants.
Allo-AVA 数据集概述
基本信息
- 许可证: cc
- 语言:
- 英语 (en)
- 标签:
- 代码 (code)
- 数据规模:
- 大于1TB (n>1T)




