遇见数据集

FINAL-Bench/World-Model

收藏
Hugging Face2026-05-15 更新2026-04-05 收录
官方服务:

资源简介:

World Model Bench(WM Bench)是首个用于评估世界模型和具身AI系统认知能力的基准数据集。它包含100个场景,通过三个支柱(感知、认知、具身)和10个类别进行综合评分,总分1000分,旨在衡量模型在场景理解、预测推理、威胁类型区分响应、自主情感升级、上下文记忆利用、后威胁自适应恢复、运动情感表达、实时认知-行动性能和身体交换可扩展性等方面的能力。数据集通过JSON格式输入,无需3D环境,仅通过文本I/O进行评估,适用于API、性能指标和实时演示等多种参与轨道。

World Model Bench (WM Bench) is the first benchmark dataset for evaluating the cognitive capabilities of world models and embodied AI systems. It contains 100 scenarios, and implements comprehensive scoring across three pillars (perception, cognition, and embodiment) and 10 categories, with a total score of 1000 points. This benchmark aims to assess models' abilities in scene understanding, predictive reasoning, threat-type discrimination and response, autonomous emotional escalation, contextual memory utilization, post-threat adaptive recovery, motor emotional expression, real-time cognitive-action performance, and physical exchange scalability. The dataset accepts JSON-formatted input, does not require a 3D environment, and conducts evaluations solely via text I/O, supporting multiple participation tracks including APIs, performance metrics, and real-time demonstrations.

提供机构:
FINAL-Bench
二维码
社区交流群
二维码
科研交流群
商业服务