REFED
收藏资源简介:
REFED数据集是通过参考级反馈机制生成的,包含了10000条指令-响应对。该数据集由伊利诺伊大学厄巴纳-香槟分校的研究团队创建,旨在通过利用高质量的参考样本中的反馈来指导新数据的合成,进而提高数据合成的质量标准。数据集的构建基于LIMA训练数据集,利用GPT-4o mini模型进行数据合成。REFED数据集可应用于指令微调任务,通过该数据集微调的模型在AlpacaEval 2.0和Arena-Hard基准测试中表现出色。
The REFED dataset is generated via a reference-level feedback mechanism, containing 10,000 instruction-response pairs. This dataset was created by the research team from the University of Illinois Urbana-Champaign, aiming to guide the synthesis of new data by leveraging feedback from high-quality reference samples, thereby improving the quality standards of data synthesis. The construction of the dataset is based on the LIMA training dataset, and the GPT-4o mini model is utilized for data synthesis. The REFED dataset can be applied to instruction fine-tuning tasks, and models fine-tuned with this dataset have achieved excellent performance on the AlpacaEval 2.0 and Arena-Hard benchmarks.

- 1Beyond Sample-Level Feedback: Using Reference-Level Feedback to Guide Data Synthesis伊利诺伊大学厄巴纳-香槟分校 · 2025年



