Havoc999/The-Claude-Dataset
收藏资源简介:
该数据集是一个大规模、经过清理和统一处理的集合,包含624,252个推理轨迹,这些轨迹蒸馏自Claude 3.5 Sonnet和Opus(版本4.6/4.7)。它专门设计用于训练小型高性能模型(例如nanowhale-100m),这些模型采用多头部潜在注意力(MLA)或混合专家(MoE)架构,其中高密度思维链数据至关重要。数据集通过聚合五个高质量蒸馏源(包括纯推理轨迹、大规模指令遵循和过滤推理样本)构建,并解决了模式规范化、箭头对齐和清理等技术壁垒,确保数据格式统一为OpenAI/HuggingFace消息格式。数据以Parquet格式提供,总行数为624,252,优化用于MLA/MoE架构的训练。
This dataset is a massive, cleaned, and unified collection of 624,252 reasoning traces distilled from Claude 3.5 Sonnet and Opus (versions 4.6/4.7). It was specifically engineered to train small, high-performance models (like nanowhale-100m) that utilize Multi-Head Latent Attention (MLA) or Mixture of Experts (MoE) architectures, where high-density Chain-of-Thought (CoT) data is critical. The dataset aggregates five high-quality distillation sources, including pure reasoning traces, large-scale instruction following, and filtered reasoning samples, and addresses technical barriers such as schema normalization, arrow alignment, and cleaning to ensure a unified OpenAI/HuggingFace Message format. It is provided in Parquet format with 624,252 rows and optimized for training MLA/MoE architectures.



