ListenerX
收藏资源简介:
ListenerX是一个大规模的3D对话数据集,包含了超过140万个有效帧,用于多模态响应交互。该数据集由北京邮电大学、香港科技大学和中国科学院自动化研究所共同创建。数据集内容涵盖了高质量的长期对话视频片段,包括对话双方的头部动作、说话人的音频、详细的文本描述以及情感强度标签。ListenerX的创建过程采用了先进的3D面部角色估计方法和面部表情分析技术,确保了数据集的准确性和多样性。该数据集旨在解决目前多模态响应交互中数据集稀缺的问题,为人类交互分析、说话人脸生成和响应交互建模等下游任务提供了基础。
ListenerX is a large-scale 3D dialogue dataset containing over 1.4 million valid frames for multimodal response interaction. This dataset was jointly created by Beijing University of Posts and Telecommunications, The Hong Kong University of Science and Technology, and the Institute of Automation of the Chinese Academy of Sciences. The dataset covers high-quality long-duration conversational video clips, including head movements of both conversational parties, audio of the speakers, detailed textual descriptions, and emotion intensity labels. The construction of ListenerX adopts advanced 3D facial character estimation methods and facial expression analysis technologies to ensure the accuracy and diversity of the dataset. This dataset aims to address the current scarcity of datasets in multimodal response interaction, providing a foundational resource for downstream tasks such as human interaction analysis, talking face generation, and response interaction modeling.

- 1VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction北京邮电大学,中国北京市;香港科技大学,中国香港;中国科学院自动化研究所,中国北京市 · 2025年



