jescy525/archon-sft-v1-code
收藏资源简介:
archon-sft-v1-code数据集是AETHER家族的监督微调(SFT)数据集,专注于代码相关任务。数据格式为JSONL ChatML消息,包含角色(如系统、用户、助手)、内容、任务类型(如函数调用、代码、推理链等)、来源数据集标识、语言和系统来源字段。应用了MinHash去重技术(阈值0.85)。数据集生成于2026年5月25日,主要来源于ise-uiuc/Magicoder-OSS-Instruct-75K(占85.7%)和m-a-p/CodeFeedback-Filtered-Instruction(占14.3%)。任务类型分布为:代码任务占51%,数学符号任务(从包含公式的代码中自动标记)占48%,protobuf_grpc任务占1.4%。尽管元数据标签显示支持英语和法语,但实际检测语言为100%英语。该数据集专用于代码指令跟随的监督微调,代表实际执行版本,与初始设计可能有所不同。
The archon-sft-v1-code dataset is an AETHER family Supervised Fine-Tuning (SFT) dataset focused on code-related tasks. The data is in JSONL ChatML message format, including fields such as role (e.g., system, user, assistant), content, task type (e.g., function_calling, code, reasoning_cot), source dataset identifier, language, and system source. MinHash deduplication is applied with a threshold of 0.85. Generated on May 25, 2026, the dataset is primarily sourced from ise-uiuc/Magicoder-OSS-Instruct-75K (85.7%) and m-a-p/CodeFeedback-Filtered-Instruction (14.3%). The task type distribution is: 51% code, 48% math_symbolic (auto-tagged from code with formulas), and 1.4% protobuf_grpc. Although metadata tags indicate support for English and French, the actual detected language is 100% English. This dataset is designed for code instruction-following supervised fine-tuning and represents the executed version, which may differ from the initial design.




