Hummer
收藏资源简介:
Hummer是一个创新的成对偏好数据集,旨在减少对齐目标之间的竞争。该数据集基于UltraFeedback构建,并通过GPT-4的AI反馈进行了增强,成为首个旨在减少对齐目标间竞争的偏好数据集。Hummer通过三阶段过程构建:偏好与目标标注、对齐目标细化及数据集分割。数据集的应用领域包括特定领域的进一步微调和减少对攻击的脆弱性,旨在通过优先考虑某些对齐目标而不牺牲其他目标的性能来解决特定问题。
Hummer is an innovative pairwise preference dataset developed to mitigate competition among alignment targets. Built upon UltraFeedback and augmented with AI feedback from GPT-4, it represents the first preference dataset specifically designed to address this competition issue. Hummer is constructed through a three-stage pipeline: preference and target annotation, alignment target refinement, and dataset splitting. Its potential applications include further domain-specific fine-tuning and reducing vulnerability to adversarial attacks, with the goal of solving targeted problems by prioritizing certain alignment targets while maintaining the performance of other targets.

- 1Hummer: Towards Limited Competitive Preference Dataset麦吉尔大学, 北京大学, 蚂蚁集团 · 2024年



