miria0/EduFeedback
收藏资源简介:
EduFeedback是一个在教育教学场景中合成的多轮对话偏好数据集,由GPT-4o生成,模拟导师(agent1)和学生(agent2)之间的对话,涵盖11个科学和哲学主题,并包含不同的情绪/个性和采样温度。该数据集在COALA论文(Feng & Pilanci, ICML 2026)中引入,并伴随一种名为交替填充策略的新方法,用于直接从多轮对话中提取高质量的(提示、选中、拒绝)偏好训练三元组,无需外部奖励模型或额外生成调用。数据集提供三种配置:原始对话、每个对话单对偏好对以及通过交替填充策略生成的每个对话多对偏好对,适用于偏好微调任务,如DPO、ORPO等。
EduFeedback is a synthetically generated, multi-turn conversational preference dataset in an educational tutoring setting, introduced in the COALA paper (Feng & Pilanci, ICML 2026). It is produced by GPT-4o acting as two agents—a tutor (agent1) and a student (agent2)—across eleven topics in science and philosophy, with varying mood/personality prompts and sampling temperatures. The dataset is released together with the Alternating Population Strategy, a novel method for extracting high-quality (prompt, chosen, rejected) preference training triplets directly from multi-turn conversations without external reward models or extra generation calls. It includes three configurations: raw conversations, single preference pair per conversation, and multiple preference pairs per conversation generated via the Alternating Population Strategy, suitable for preference fine-tuning tasks such as DPO, ORPO, etc.




