Kwai-Klear/GoLongRL
收藏官方服务:
资源简介:
该数据集是GoLongRL的强化学习训练数据集,旨在提升语言模型的长上下文能力。它总共包含23,000个训练样本,并涵盖9种奖励函数类型。
This dataset is the GoLongRL reinforcement learning training dataset, purpose-built to enhance the long-context capabilities of language models. It contains a total of 23,000 training samples and covers 9 types of reward functions.
提供机构:
Kwai-Klear


