We propose ScheduleNet, a RL-based real-time scheduler, that can solve various types of multi-agent scheduling problems. We formulate these problems as a semi-MDP with episodic reward (makespan) and l
这些数据基于anthropic的论文《Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback》开源,旨在为后续的RLHF训练训练偏好(或奖励)模型。数据集来源于beyond/rlhf-reward-single-round-trans_chinese,并使用OpenCC进行简