官方服务:
资源简介:
Top episode replays for 2026-05-10.
2026年5月10日的顶级剧集回放
应用场景:
创建时间:
2026-05-19
相关数据集
disserji1/OpenThoughts-Agent-v1-RL
OpenThoughts-Agent-v1-RL 是一个精心策划的强化学习数据集,包含约720个任务,每个任务都有指令、环境和验证器,用于代理训练。该数据集用于训练 OpenThinker-Agent-v1 模型的强化学习阶段,任务来源于 nl2bash 验证数据集,并经过三阶段过滤管道以确保质量。数据集还包括监督微调(SFT)轨迹数据集 OpenThoughts-Agent-v1-SFT,包含约
Hugging Face2025-12-16 更新150
skandermoalla/qrpo-paper-mistral-nosft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm
--- license: mit tags: - reinforcement-learning - alignment - qrpo --- # qrpo-paper-mistral-nosft-magpieair-armorm-temp1-ref50-offpolicy2best-armorm Dataset with reference completions and rewards fo
Hugging Face2025-12-08 更新70
skandermoalla/qrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offpolicy2random-armorm
--- license: mit tags: - reinforcement-learning - alignment - qrpo --- # qrpo-paper-mistral-sft-magpieair-armorm-temp1-ref50-offpolicy2random-armorm Dataset with reference completions and rewards fo
Hugging Face2025-12-08 更新50
KHuss/hh-rlhf-formatted
该数据集是Anthropic的hh-rlhf数据集的重新格式化版本,原始数据集可在HuggingFace上获取。RLHF数据集通常定义为三元组D={(x,y_1,y_2)_n}的集合,其中x是提示,y_1是选择的响应,y_2是被拒绝的响应。原始数据集提供了两列,chosen=x+y_1和rejected=x+y_2。为了适应不同的RLHF设置,本数据集保留了原始的chosen和rejected列,
Hugging Face2024-11-10 更新80
NN DynaQ with prioritized sweeping and multiple predecessors
Figures illustrating the model, the task and the results of an neural network implementation of a DynaQ reinforcement learning algorithm with prioritized sweeping, in a case where a (state,action) cou
DataCite Commons2020-08-30 更新50



