遇见数据集

Stable-Alignment

收藏
OpenXLab2026-04-18 收录
官方服务:

资源简介:

This is the official repo for the Stable Alignment project. We aim to provide a RLHF alternative which is superior in alignment performance, highly-efficient in data learning, and easy to deploy in scaled-up settings. Instead of training an extra reward model that can be gamed during optimization, we directly train on the recorded interaction data in simulated social games. We find high-quality data + reliable algorithm is the secret recipe for stable alignment learning.

提供机构:
OpenDataLab
创建时间:
2024-04-30
二维码
社区交流群
二维码
科研交流群
商业服务