Omni-RewardData
收藏资源简介:
Omni-RewardData是一个包含多种模态偏好的大规模多模态偏好数据集,旨在训练能够处理所有模态的通用多模态奖励模型。数据集由248K个通用偏好对和69K个包含自由形式偏好描述的指令调整对组成,覆盖了文本、图像、视频、音频和3D等五个模态。该数据集的构建旨在解决现有奖励模型在模态不平衡和偏好刚性方面的挑战,从而提高模型的泛化能力,更好地适应不同的用户偏好。
Omni-RewardData is a large-scale multi-modal preference dataset encompassing diverse modal preferences, aiming to train general multi-modal reward models capable of handling all modalities. The dataset consists of 248K general preference pairs and 69K instruction tuning pairs with free-form preference descriptions, covering five modalities including text, image, video, audio, and 3D. The construction of this dataset is designed to address the challenges of modal imbalance and preference rigidity in existing reward models, thereby enhancing the generalization ability of models and better adapting to diverse user preferences.




