Uni-DPO
收藏资源简介:
Uni-DPO数据集是一个用于大型语言模型(LLMs)动态偏好优化的多模态数据集,涵盖文本理解、数学推理和多模态理解三个关键领域。该数据集旨在通过联合考虑偏好对的内在质量和模型学习动态,实现更有效和稳健的偏好学习。数据集包含从高质量来源生成的偏好对,如HuggingFaceH4/ultrafeedback_binarized和RLHFlow/numia_prompt_dpo1,并通过特定的脚本和流程进行生成和标注。Uni-DPO的主要优势包括质量感知、动态感知和轻量级统一框架,能够自适应地优先处理高质量偏好对并减轻过拟合。该数据集适用于文本到文本、多模态LLM、偏好学习和RLHF等任务,规模在10万到100万样本之间。
The Uni-DPO dataset is a multimodal dataset for dynamic preference optimization of Large Language Models (LLMs), covering three core areas: text understanding, mathematical reasoning, and multimodal understanding. This dataset aims to achieve more efficient and robust preference learning by jointly considering the intrinsic quality of preference pairs and the model's learning dynamics. The dataset contains preference pairs generated from high-quality sources, such as HuggingFaceH4/ultrafeedback_binarized and RLHFlow/numia_prompt_dpo1, and is generated and annotated via specific scripts and workflows. The core advantages of Uni-DPO include quality-awareness, dynamics-awareness, and a lightweight unified framework, which can adaptively prioritize high-quality preference pairs and mitigate overfitting. This dataset is applicable to tasks such as text-to-text, multimodal LLM, preference learning, and RLHF, with a scale ranging from 100,000 to 1,000,000 samples.



