遇见数据集

MoAA-SFT

收藏
魔搭社区2026-04-28 更新2026-07-15 收录
官方服务:

资源简介:

This is the SFT data of our MoAA method described in this [paper](https://arxiv.org/abs/2505.03059). We subsample from two widely-used open-source instruction tuning datasets: UltraFeedback and UltraChat. Our subsampling strategy involves utilizing the entire UltraFeedback dataset and randomly selecting 5,000 samples from UltraChat. We use MoA to generate responses. The proposers used in our study are WizardLM-2-8x22b, Gemma-2-7b-it, Qwen-2-72b-Instruct, and Llama-3.1-70b-Instruct, while Qwen-1.5-110b-Instruct serves as the aggregator. ## Citation ``` @article{wang2025improving, title = {Improving Model Alignment Through Collective Intelligence of Open-Source LLMS}, author = {Junlin Wang and Roy Xie and Shang Zhu and Jue Wang and Ben Athiwaratkun and Bhuwan Dhingra and Shuaiwen Leon Song and Ce Zhang and James Zou}, year = {2025}, journal = {arXiv preprint arXiv: 2505.03059} } ```

提供机构:
maas
创建时间:
2025-11-18
二维码
社区交流群
二维码
科研交流群
商业服务