gemma-vs-gemma-preferences
收藏资源简介:
该数据集包含使用[anakin87/gemma-2-2b-ita-sft](https://huggingface.co/anakin87/gemma-2-2b-ita-sft)生成的策略内收集的偏好数据。数据集主要用于意大利语的文本生成任务,特别是用于偏好优化(Preference Optimization)和直接偏好优化(DPO)的研究。数据集包含通过特定模型生成的响应,并使用Llama模型进行评估和排名。虽然该数据集对教学目的有价值,但不建议用于偏好调优的训练,因为训练将是离策略的(off-policy),并且数据集是由一个小型意大利语模型生成的。
This dataset contains on-policy preference data collected using the model [anakin87/gemma-2-2b-ita-sft](https://huggingface.co/anakin87/gemma-2-2b-ita-sft). It is primarily intended for Italian text generation tasks, especially research on preference optimization and Direct Preference Optimization (DPO). The dataset includes responses generated by the specified model, which are evaluated and ranked using Llama models. While this dataset has value for educational purposes, it is not recommended for preference tuning training, as such training would be off-policy, and the dataset was generated by a small Italian language model.




