Xuhui/sim-posttrain
收藏资源简介:
HUMANUAL后训练数据是一个用于用户模拟任务的数据集,包含多种配置,涵盖了新闻、政治、观点、书评、聊天和电子邮件回复等多个领域。数据集还包括专门的评估集,如UserLM评估、错误评估、Social-R1评估、SocSci210和HumanLLM项目选择,每个评估集都有特定的用例和评估标准。数据集的结构详细,每个配置的字段和用途都有明确说明。该数据集旨在用于Harmony中的RL后训练和评估,并提供了如何为不同任务计算奖励的具体指导。
The HUMANUAL Posttraining Data dataset is designed for user simulation tasks, featuring multiple configurations across various domains such as news, politics, opinion, book reviews, chat, and email responses. It also includes specialized evaluation sets like UserLM Eval, Mistakes Eval, Social-R1 Eval, SocSci210, and HumanLLM Item Selection, each with specific use cases and evaluation metrics. The datasets schema is detailed, with clear explanations of each fields structure and purpose. It is intended for use in RL posttraining and evaluation within Harmony, providing specific instructions on reward computation for different tasks.




