遇见数据集

kaushik-harsh-99/Uncensored-SFT-v1

收藏
Hugging Face2026-05-07 更新2026-05-31 收录
官方服务:

资源简介:

这是一个高质量、未经审查的指令数据集,专为英语设计,通过从Hugging Face收集多个公开指令数据集(包括指令遵循、聊天、问答、角色扮演、编码等类型)并经过多阶段处理流程构建而成。处理流程包括格式标准化、OCR和垃圾清理、英语过滤、精确去重、问题去重、最佳响应选择和低质量过滤,最终包含约721,000行数据,采用JSONL格式,每行包含input和output字段。该数据集适用于监督微调、指令调优、聊天模型训练、对齐实验和未经审查的助手研究等任务,尤其适合小型模型的能力提升和减少过度拒绝行为。

This is a high-quality, uncensored instruction dataset designed for English, built by collecting multiple public instruction datasets from Hugging Face (including instruction-following, chat, QA, roleplay, coding, and other types) and processed through a multi-stage pipeline. The pipeline includes format normalization, OCR and junk cleaning, English-only filtering, exact deduplication, question deduplication, best response selection, and low-quality filtering, resulting in approximately 721,000 rows in JSONL format, with each row containing input and output fields. The dataset is intended for supervised fine-tuning, instruction tuning, chat model training, alignment experimentation, uncensored assistant research, and other tasks, particularly suitable for improving smaller models capabilities and reducing over-refusal behavior.

提供机构:
kaushik-harsh-99
二维码
社区交流群
二维码
科研交流群
商业服务