CAD-bench/cad-bench-ed-2026-anonymous-full
收藏资源简介:
该数据集包含完整的匿名CAD-bench评审工件,包括公共任务负载、报告结果JSON、源代码存档和用于论文表格的运行证明报告工件。每个任务目录包含多个文件,如prompt.txt(自然语言基准提示)、task.toml(任务元数据、难度、评估者名称和期望值)、gold.py(用于验证和媒体生成的参考Build123D解决方案)以及可选的夹具(如STEP文件或Blender模拟脚本)。数据集旨在与CAD-bench运行时一起使用,以评估CAD代码生成或代理CAD系统。数据来源为合成的CAD提示和基准元数据,不包含个人数据或人类受试记录。许可方面,创作的基准代码、提示、任务元数据和参考程序在MIT许可下发布。数据集包含17个任务,其中一些简单的几何任务接近当前模型的解决范围,而功能组装任务仍然困难。
This dataset contains the full anonymous CAD-bench reviewer artifact. It includes the public task payloads, reported result JSON, source archive, and provenance report artifacts for the runs used by the paper tables. Each task directory includes files such as prompt.txt (the natural-language benchmark prompt), task.toml (task metadata, difficulty, evaluator name, and expected values), gold.py (a reference Build123D solution used for validation and media generation), and optional fixtures like STEP files or Blender simulation scripts. The dataset is intended for use with the CAD-bench runtime to evaluate CAD code-generation or agentic CAD systems. The data provenance consists of synthetic CAD prompts and benchmark metadata authored for this benchmark, with no personal data or human-subject records. Licensing-wise, authored benchmark code, prompts, task metadata, and reference programs are released under MIT. The dataset has 17 tasks, with some simple geometry tasks close to being solved by current models, while functional assembly tasks remain difficult.




