jescy525/archon-sft-v1-conversation
收藏资源简介:
archon-sft-v1-conversation是一个AETHER家族的有监督微调数据集,属于对话组。该数据集采用JSONL ChatML消息格式,应用了任务类型标签和MinHash去重技术(阈值为0.85)。数据模式包括消息(包含角色如系统、用户或助手及内容)、任务类型(如函数调用、代码、推理链等)、来源数据集、语言(主要为英语)和系统来源。数据集生成于2026年5月25日,通过prepare_sft.py管道处理,实际数据来源100%来自HuggingFaceH4/ultrachat_200k。任务类型分布为:78%通用对话、16%推理链、4%数学符号(自动标记)和1.6%协议缓冲区/gRPC。尽管元数据标签包含英语和法语,但采样显示语言为100%英语。这是一个通用对话有监督微调数据集,代表实际数据而非初始设计规划。
archon-sft-v1-conversation is an AETHER family SFT dataset — group conversation. The format uses JSONL ChatML messages with task_type tagging and MinHash deduplication applied (threshold 0.85). Schema includes messages (with roles like system, user, or assistant and content), task_type (e.g., function_calling, code, reasoning_cot), source_ds, lang (primarily English), and system_source. Generated by the prepare_sft.py pipeline on 2026-05-25, with real sources sampled as 100% from HuggingFaceH4/ultrachat_200k. Task type distribution: 78% general_conv, 16% reasoning_cot, 4% math_symbolic (auto-tagged), and 1.6% protobuf_grpc. Language reality: sampled rows show 100% English despite metadata tags including English and French. This is a general conversational SFT dataset representing the actual data, not the initial design plan.




