awesome-text-to-motion
收藏资源简介:
这是一个关于文本驱动人体运动生成的数据集资源集合,专注于单人场景且不涉及人-物/场景交互。该合集收录了该领域相关的多个数据集,并提供了数据集的统计图表和交互式可视化,旨在为研究人员和开发者提供全面的数据集索引和整理。
This is a collection of dataset resources focused on text-driven human motion generation, targeting single-person scenarios with no human-object or scene interactions involved. This compilation includes multiple datasets relevant to the field, alongside statistical charts and interactive visualizations for each dataset, with the goal of providing researchers and developers with comprehensive dataset indexing and organization.
awesome-text-to-motion 数据集详情
项目概述
这是一个专注于文本驱动的人体运动生成(Text-driven Human Motion Generation)的资源集合,涵盖综述、数据集和模型,聚焦于单人场景(不涉及人与物体或场景的交互)。项目提供了交互式可视化图表和统计数据的在线页面。
综述(Surveys)
共收录 4 篇综述论文:
| 论文标题 | 发表信息 | 链接 |
|---|---|---|
| Motion Generation: A Survey of Generative Approaches and Benchmarks | arXiv(2025) | https://arxiv.org/abs/2507.05419 |
| Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation | arXiv(2025) | https://arxiv.org/abs/2506.03191 |
| Text-driven Motion Generation: Overview, Challenges and Directions | arXiv(2025) | https://arxiv.org/abs/2505.09379 |
| Human Motion Generation: A Survey | TPAMI(2023) | https://arxiv.org/abs/2307.10894 |
数据集(Datasets)
共收录 26 个数据集:
2025年发布
- UniMo4D (arXiv 2025): X-MoGen: Unified Motion Generation across Humans and Animals
- FineMotion (arXiv 2025): 细粒度时空注释的数据集和基准,用于细粒度运动生成和编辑
- SnapMoGen (arXiv 2025): 从表达性文本生成人体运动
- MotionMillion (ICCV 2025): 百万级数据驱动的零样本运动生成
- HumanAttr (arXiv 2025): 生成属性感知的人体运动
- GBC-100K (arXiv 2025): 从运动到行为的分层建模
- STANCE (CVPR 2025): 动态运动融合用于多功能运动编辑
- PerMo (CVPR 2025): 个性化文本到运动生成
- TMD (arXiv 2025): 任意到运动生成
- Motion-X++ (arXiv 2025): 大规模多模态3D全身人体运动数据集
2024年发布
- MotionFix (SIGGRAPH Asia 2024): 文本驱动的3D人体运动编辑
- HumanML3D-Extend (arXiv 2024): 长文本指令驱动的无限运动生成
- MotionPercept (ICLR 2025): 对齐人体运动生成与人类感知
- PaM (arXiv 2024): 通过AI反馈驱动的直接偏好优化
- HumanML3D++ (ICCV 2025): 任意文本到运动生成
- MotionVerse (arXiv 2024): 统一多模态运动生成的大运动模型
- RICH-CAT (arXiv 2024): 接触感知的文本描述人体运动生成
- FineHumanML3D (LREC-COLING 2024): 从细粒度文本描述生成运动
- BlindWays (NeurIPS 2024): 文本到盲人运动
- LaViMo (ECCV 2024): 开放词汇描述的双向3D人体运动生成
- Inter-MT2 (ICCV 2025): 人类交互中运动推理与生成的统一框架
- MotionLib (ICML 2025): 百万级人体运动的大运动模型扩展
- HumanML3D-synthesis (MM 2024): 文本驱动人体运动生成的性能评估库
- Limb-ET2M (MM 2024): 通过LLM引导的肢体级情感操控实现情感增强的文本到运动生成
2023年及之前
- Motion-X (NeurIPS 2023): 大规模3D表达性全身人体运动数据集
- HumanLong3D (AAAI 2024): 自回归运动扩散
- HuMMan-MoGen (NeurIPS 2023): 细粒度时空运动生成与编辑
- HumanML3D (CVPR 2022): 从文本生成多样且自然的3D人体运动
- KIT (Big Data 2016): KIT运动语言数据集
模型(Models)
共收录 52 个模型,涵盖多种技术路线:
代表性模型类别
| 模型名称 | 发表信息 | 关键技术 |
|---|---|---|
| X-MoGen | arXiv 2025 | 跨人类与动物的统一运动生成 |
| ReMoMask | arXiv 2025 | 检索增强的掩码运动生成 |
| SASI | SIGGRAPH 2025 | 语义一致的文本到运动与无监督风格 |
| MoMask++ | arXiv 2025 | 表达性文本的运动生成 |
| GotoZero | ICCV 2025 | 零样本运动生成 |
| MotionGPT3 | arXiv 2025 | 人体运动作为第二模态 |
| Motion-R1 | arXiv 2025 | 思维链推理与强化学习 |
| MOGO | arXiv 2025 | 残差量化分层因果Transformer |
| FlowMotion | arXiv 2025 | 目标预测条件流匹配 |
| MixerMDM | CVPR 2025 | 可学习的运动扩散模型组合 |
| SALAD | CVPR 2025 | 骨架感知的潜扩散模型 |
| MotionLab | ICCV 2025 | 统一运动生成与编辑 |
| PackDiT | arXiv 2025 | 联合人体运动和文本生成 |
其他模型(按时间排序)
ACMDM、MoMADiff、ReAlign、GENMO、DSDFM、UniPhys、Shape-Move、MG-MotionLLM、ReMoGPT、UniTMGE、MotionReFit、LoRA-MDM、HMU、SimMotionEdit、MotionStreamer、GenM³、Kinesis、sMDM、PMG、PersonaBooth、MotionAnything、BioVAE、MoMug、Fg-T2M++、CASIM、SPORT、MotionPCM、Free-T2M、FlexMotion、MMDM、EgoLM、MoGenTS 等。
技术特点
- 聚焦范围: 单一人体运动生成,不涉及人-物体或人-场景交互
- 核心任务: 文本到运动生成(Text-to-Motion Generation)
- 主流技术: 扩散模型、自回归模型、检索增强、强化学习、大语言模型
- 研究方向: 零样本生成、细粒度控制、个性化生成、情感增强、物理一致性等




