Streamer
收藏资源简介:
Streamer数据集是一个包含中文语音和显式语义手势的3D人体运动高质量数据集,主要用于研究直播场景下的手语合成。该数据集包含了一系列预定义的手势(如数字和方向),以及相应的语音音频。Streamer数据集的创建旨在解决现有数据集中缺乏显式语义手势的问题,并通过混合模态扩散Transformer和级联同步检索增强生成技术,实现灵活可靠的人体手势合成。该数据集的应用领域包括电影制作、游戏设计、机器人技术和数字人类创建等。
The Streamer Dataset is a high-quality 3D human motion dataset encompassing Chinese speech and explicit semantic gestures, primarily intended for research on sign language synthesis in live streaming scenarios. This dataset includes a series of predefined gestures (e.g., numbers and directions) along with corresponding speech audio. Developed to address the shortage of explicit semantic gestures in existing datasets, the Streamer Dataset enables flexible and reliable human gesture synthesis via multimodal diffusion Transformer and cascaded synchronous retrieval-augmented generation technologies. Its applicable domains include film production, game design, robotics, and digital human creation, among others.
GestureHYDRA 数据集概述
基本信息
- 数据集名称: GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
- 发表会议: ICCV 2025
- 作者机构:
- 中国科学技术大学
- 百度VIS团队
- 贡献者: Quanwei Yang, Luying Huang, Kaisiyuan Wang等
- 相关资源:
- 论文
- arXiv
- 视频
- 代码
- 数据
数据集描述
- 主要内容: 包含3D人体运动数据,特别是一组具有明确语义的手势,常用于直播主播。
- 特点:
- 高质量
- 语义明确的手势
- 用于语音伴随手势生成研究
研究方法
- 系统架构:
- 基于混合模态扩散变换器架构
- 包含新颖设计的运动风格注入变换器层
- 关键技术:
- 级联检索增强生成策略
- 自适应音频-手势同步机制
- 语义手势库
实验结果
- 比较数据集:
- TalkSHOW数据集
- Streamer数据集
- 生成结果展示:
- 自适应关键帧注入结果
- 手势编辑结果
- 基于合成手势的视频生成
引用格式
bibtex @inproceedings{yang2025GestureHYDRA, author = {Quanwei Yang, Luying Huang, Kaisiyuan Wang, Jiazhi Guan, Shengyi He, Fengguo Li, Lingyun Yu, Yingying Li, Haocheng Feng, Hang Zhou, Hongtao Xie.}, title = {GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation}, booktitle = {ICCV}, year = {2025}, }
许可信息
- 许可证类型: Creative Commons Attribution-ShareAlike 4.0 International License




