DyConv
收藏资源简介:
DyConv是由字节跳动公司创建的一个大型双人对话数据集,旨在支持音频驱动的交互式头部生成研究。该数据集包含从互联网收集的非脚本视频对话,涵盖了多样化的背景和广泛的讨论主题及情感,真实反映了现实生活中的交流场景。数据集的创建过程涉及从大量真实对话视频中提取面部交流行为,并将其编码为低维运动潜在空间。DyConv的应用领域主要集中在构建能够自然切换听讲状态的对话代理,解决现有模型在双人对话中角色分配和切换不自然的问题。
DyConv is a large-scale two-person dialogue dataset created by ByteDance, aimed at supporting research on audio-driven interactive head generation. This dataset includes unscripted video dialogues collected from the Internet, covering diverse backgrounds, a wide range of discussion topics and emotions, and faithfully reflects real-life communication scenarios. The dataset construction process entails extracting facial communicative behaviors from a large number of real dialogue videos and encoding these behaviors into a low-dimensional motion latent space. The primary application scope of DyConv is to build dialogue agents that can seamlessly switch between listening and speaking states, addressing the problem of unnatural role assignment and switching in existing two-person dialogue models.




