processed_dolphin_sft_v0.1_preference
收藏官方服务:
资源简介:
Processed to have reasoning accepted for ORPO fine tuning The preference dataset was generated using Mistral-Instruct-v0.1 finetuned on a GPT-4 subset of the Dolphin dataset! Generated responses are labeled as rejected, GPT-4 responses (original Dolphin data) are labeled as accepted. The motivation was to test out the SPIN paper finetuning methodology. [Link to the dataset](https://huggingface.co/datasets/reciperesearch/dolphin-sft-v0.1-preference)
提供机构:
maas创建时间:
2026-07-28



