WORLDCODER-BENCH
收藏资源简介:
WORLDCODER-BENCH是由中国科学院自动化研究所与华为诺亚方舟实验室联合创建的首个专注于物理基础3D世界合成的基准数据集,旨在评估大型语言模型生成可执行、交互式Three.js程序的能力。该数据集包含2,026个专家精心策划的任务,覆盖模拟、渲染和应用三大宏观类别及15个细粒度领域,数据来源于人工种子创建与LLM辅助扩展,并经过严格过滤与验证。数据集创建过程采用四阶段流水线,包括专家种子策划、扩展过滤、运行时验证及反污染处理,确保了任务的高质量与多样性。其主要应用于评估和提升AI模型在物理正确性、资产集成和状态同步方面的能力,旨在解决生成式3D世界中行为正确性难以从外部观测的挑战,推动交互式3D内容开发的自动化进程。
WORLDCODER-BENCH is the first benchmark dataset dedicated to physically grounded 3D world synthesis, jointly created by the Institute of Automation of the Chinese Academy of Sciences and Huawei Noah's Ark Lab. It aims to evaluate the capability of large language models (LLMs) to generate executable and interactive Three.js programs. This dataset contains 2,026 expert-curated tasks, covering three macro categories: simulation, rendering, and application, as well as 15 fine-grained domains. The data is sourced from manual seed creation and LLM-aided expansion, and has undergone rigorous filtering and validation. The dataset creation process adopts a four-stage pipeline, including expert seed curation, expansion filtering, runtime verification, and anti-contamination processing, to ensure the high quality and diversity of the tasks. Its main applications are to evaluate and enhance the capabilities of AI models in terms of physical correctness, asset integration, and state synchronization. It aims to address the challenge that the behavioral correctness of generative 3D worlds is difficult to observe externally, and promote the automation process of interactive 3D content development.





