zimplex/genesis-hr-bench-dp3-per-task-progressive-20260529
收藏资源简介:
Genesis HR Bench — DP3每任务渐进式评估数据集(版本per_task_20260529_051038)是一个用于机器人策略评估的数据集,包含44个任务在DP3(3D Diffusion Policy)训练和评估管道中的渐进式上传结果。每个任务在单个H200 GPU上并行训练,训练时间至少4小时后,系统会自动为每个新检查点提交评估任务,数据集收集这些评估结果。数据集结构按任务组织,每个任务文件夹下包含按epoch和测试平均分数命名的子目录,其中存储评估结果文件(如results.json、eval.log、manifest.txt)、视频文件(前3个episode的mp4视频)和动作轨迹数据。数据集的生成涉及训练脚本、后台监控程序、评估脚本和上传脚本。用户可以通过HuggingFace Hub下载数据集,并筛选特定任务。数据集主要用于分析策略质量与训练时间的关系,但部分任务存在模拟问题(如缺少资产或Genesis模拟错误)。
Genesis HR Bench — DP3 Per-Task Progressive Evaluation Dataset (version per_task_20260529_051038) is a dataset for robotic policy evaluation. It contains progressive upload results of 44 tasks through the DP3 (3D Diffusion Policy) training and evaluation pipeline. Each task is trained in parallel on a single H200 GPU. After at least 4 hours of training, the system automatically submits evaluation tasks for each new checkpoint, and the dataset collects these evaluation results. The dataset is organized by task: each task folder contains subdirectories named after epoch and test average score, which store evaluation result files (e.g., results.json, eval.log, manifest.txt), video files (mp4 videos of the first 3 episodes), and motion trajectory data. The generation of this dataset involves training scripts, background monitoring programs, evaluation scripts, and upload scripts. Users can download the dataset via the Hugging Face Hub and filter for specific tasks. This dataset is primarily used to analyze the relationship between policy quality and training time, though some tasks have simulation issues such as missing assets or Genesis simulation errors.




