UltraFeedback-chinese
收藏资源简介:
UltraFeedback-Chinese是根据UltraFeedback数据集的构建方法制定的中文版本,专为训练强大的奖励模型和批评模型而设计。该数据集支持PPO(近端策略优化)和DPO(直接偏好优化)两种训练方式。数据收集自多个中文资源库,涵盖了约58k条中文指令,并对每个指令生成4个模型响应。数据集变体UltraFeedback-Chinese-Binarized专为DPO训练优化,通过设定权重对每个响应的分数进行加权,以计算得到每个响应的综合评分。实验结果表明,该数据集在提升中文语言模型表现方面具有显著效果。
UltraFeedback-Chinese is the Chinese adaptation of the UltraFeedback dataset, developed using its original construction methodology, and is specifically designed for training robust reward models and critic models. This dataset supports two training paradigms: Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO). The data is collected from multiple Chinese language corpora, covering approximately 58k Chinese instructions, with 4 model responses generated for each instruction. The dataset variant UltraFeedback-Chinese-Binarized is specially optimized for DPO training, where the scores of each response are weighted with preset weights to calculate the comprehensive score for individual responses. Experimental results demonstrate that this dataset yields significant performance improvements for Chinese language models.




