MMBench2
收藏资源简介:
MMBench2是由加州大学圣地亚哥分校研究团队构建的大规模多任务视觉世界建模数据集,旨在系统研究生成世界模型中的幻觉问题。该数据集包含65,600条轨迹,总计427小时的224×224分辨率视频(15帧/秒),覆盖210个连续控制任务,数据来源包括10个不同领域如DMControl、Meta-World等。数据集创建过程通过扩展MMBench基准,整合了真实动作标签、奖励信号及实时仿真环境。其核心应用在于探究世界模型幻觉的预测与预防机制,通过提供行为多样性数据支持模型在低覆盖区域的适应性研究。
MMBench2 is a large-scale multi-task visual world modeling dataset developed by a research team from the University of California, San Diego, aiming to systematically study hallucination issues in generative world models. This dataset contains 65,600 trajectories, totaling 427 hours of 224×224 resolution videos at 15 frames per second (fps), covering 210 continuous control tasks. Its data sources span 10 distinct domains including DMControl, Meta-World, and others. The dataset was constructed by extending the MMBench benchmark, integrating real action labels, reward signals, and real-time simulation environments. Its core application lies in exploring the prediction and prevention mechanisms of world model hallucinations, and supporting adaptive research of models in low-coverage regions by providing behaviorally diverse data.
数据集概述
数据集名称: MMBench2
数据集规模:
- 总时长:427 小时
- 任务数量:210 个任务
- 领域数量:10 个领域
- 轨迹数量:未明确给出具体数字,但每个任务包含等量的轨迹数,帧数分布不均,帧数范围从 25(ManiSkill3)到 1000(Atari)步。
任务与领域:
- 任务类型涵盖:运动控制、操作、导航、街机风格环境等。
- 涉及的领域包括:DMControl、DMControl Ext.、Meta-World、ManiSkill3、MuJoCo、Box2D、RoboDesk、OGBench、MiniArcade、Atari。
数据内容:
- 包含 ground-truth 动作(actions)、奖励(rewards)、语言指令(language instructions)。
- 每个任务都配备实时模拟器(live environment)。
数据特点:
- 帧数分布呈重尾分布,非均匀,用于研究数据覆盖问题。
开源状态:
- 完全开源(fully open-source)。
相关资源:
- 代码:提供训练和评估世界模型的代码。
- 数据集:包含 427 小时、210 个任务的数据。
- 模型:提供预训练和微调后的 350M 参数模型。

- 1Hallucination in World Models is Predictable and Preventable加州大学圣地亚哥分校 · 2026年



