遇见数据集

pperojas/dolphin

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

Dolphin数据集是一个旨在复现Microsoft Orca论文结果的文本生成数据集。它包含两个主要配置:flan1m-alpaca-uncensored(约100万条FLANv2增强数据,使用GPT-4生成完成部分)和flan5m-alpaca-uncensored(约350万条FLANv2增强数据,使用GPT-3.5生成完成部分)。数据集遵循Orca论文的子混合和系统提示分布,但进行了一些调整,例如包含了所有75k的CoT数据并去除了重复项。此外,数据集过滤了与对齐、拒绝、回避和偏见相关的内容,以生成一个未经过滤的基础模型,用户可以在其上添加个性化的对齐LoRA。数据集基于Apache-2.0许可证,允许商业或非商业使用。计划在多个基础模型上发布,包括Xgen 7b、LLaMA 13b(非商业)、MPT 30b、LLaMA 33b(非商业)、Falcon 40b和LLaMA 65b(非商业),发布模型将遵循基础模型的许可证。

The Dolphin dataset is a text-generation dataset aimed at replicating the results of Microsofts Orca paper. It consists of two main configurations: flan1m-alpaca-uncensored (approximately 1 million FLANv2 augmented data with GPT-4 completions) and flan5m-alpaca-uncensored (approximately 3.5 million FLANv2 augmented data with GPT-3.5 completions). The dataset follows the submix and system prompt distribution outlined in the Orca paper, with some exceptions, such as including all 75k CoT data and removing duplicates. Additionally, the dataset filters out instances of alignment, refusal, avoidance, and bias to produce an uncensored base model upon which personalized alignment LoRA can be layered. It is licensed under Apache-2.0 for commercial or non-commercial use. Planned releases are based on foundational models including Xgen 7b, LLaMA 13b (non-commercial), MPT 30b, LLaMA 33b (non-commercial), Falcon 40b, and LLaMA 65b (non-commercial), with releases subject to the license of the foundational model.

提供机构:
pperojas
二维码
社区交流群
二维码
科研交流群
商业服务