aditijc/snooker-testbed-phase4f2-longer-v1
收藏资源简介:
该数据集记录了斯诺克测试台第4F2阶段的训练数据,扩展了4F阶段的500k步突破到1.5M步。数据集包含15行和8列,记录了PPO算法在训练过程中的各种指标,如步数、课程阶段、平均得分、最高得分、平均击球数、平均效率、平均犯规率和回合数。训练过程中展示了持续改进,没有出现平台期,最高得分达到5.69(随机得分的2.7倍),单回合最高得分为23分。犯规率在训练过程中从99%降至93.9%。数据集还包含了生成参数和超参数的详细信息。
This dataset records the training data for Phase 4F2 of the snooker-testbed, extending the 500k-step breakthrough of Phase 4F to 1.5M steps. The dataset contains 15 rows and 8 columns, documenting various metrics during the training process of the PPO algorithm, such as steps, curriculum stage, mean score, max score, mean shots, mean efficiency, mean foul rate, and episodes. The training process demonstrated sustained improvement without plateauing, with a peak score of 5.69 (2.7x random!) and a single-episode max score of 23 points. The foul rate dropped from 99% to 93.9% during training. The dataset also includes detailed information on generation parameters and hyperparameters.



