遇见数据集

WorldModelBench

收藏
arXiv2025-09-30 收录
官方服务:

资源简介:

该数据集旨在评估视频生成模型在应用驱动领域中的世界建模能力,特别关注遵循指令和物理规律遵守方面。它包含了14种视频生成模型的评估,突出了性能指标和模型间的对比。该数据集规模达到了67,000个人类标注,任务是对视频生成模型作为世界模型的表现进行评价。

This dataset is designed to evaluate the world modeling capabilities of video generation models in application-driven domains, with special emphasis on instruction following and adherence to physical laws. It encompasses evaluations of 14 video generation models, highlighting performance metrics and inter-model comparisons. The dataset comprises 67,000 human annotations, with the core task of assessing the performance of video generation models as world models.

提供机构:
WorldModelBench team
搜集汇总
数据集介绍
WorldModelBench 数据集图片
背景与挑战
背景概述
WorldModelBench是一个用于评估视频生成模型世界建模能力的多领域基准测试,涵盖7个应用领域和56个子领域的350个提示。它通过指令遵循、常识和物理遵循(包括5个物理定律)等标准进行精细化评估,并利用67K人类标签训练了一个2B多模态评判器以提高自动化评估的准确性。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务