ulab-ai/sotopia-rl-reward-annotation
收藏资源简介:
Sotopia-RL数据集是一个用于训练社交智能代理的框架,它通过将社交互动中的剧集级反馈细化为语句级的多维奖励,提高了信用分配的准确性,并捕捉了社交行为的丰富性,从而在社会目标完成任务中取得了最先进的表现。数据集包括处理过的Sotopia-PI对话剧集、LLM生成的语句级奖励归因注释、用于训练多维奖励模型的数据以及用于GRPO训练的数据。
The Sotopia-RL dataset is a framework for training socially intelligent agents by refining episode-level feedback from social interactions into fine-grained, utterance-level, multi-dimensional rewards. This improves the accuracy of credit assignment and captures the richness of social behaviors, leading to state-of-the-art performance in social goal completion tasks. The dataset includes processed Sotopia-PI conversational episodes, LLM-generated utterance-level reward attribution annotations, data formatted for training the multi-dimensional reward model, and data for Group Reward Policy Optimization (GRPO) training.



