遇见数据集

jayshah5696/humanize-rl-tasks-v03

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

humanize-rl RL任务v03数据集包含492个强化学习任务,旨在训练和评估在特定约束下生成类人散文的模型。每一行将一个编写的指令与一个约束规范配对,该规范可由humanize-rl奖励环境(确定性检查+第一层启发式+岭回归风格评分器)使用。

492 RL tasks designed to train and evaluate models that produce human-sounding prose under specific constraints. Each row pairs an authored instruction with a constraint spec consumable by the humanize-rl reward environment (deterministic checks + Layer-1 heuristics + ridge regression style scorer).

提供机构:
jayshah5696
二维码
社区交流群
二维码
科研交流群
商业服务