遇见数据集

PGCodeLLM/ccr-bench

收藏
Hugging Face2026-02-27 更新2026-05-10 收录
官方服务:

资源简介:

# v2_human_gpt — SFT Training Data for Qwen3-Coder ## Overview This directory contains SFT training data converted to a format that works with **stock LLaMA-Factory** (no modifications) and produces **identical tokens** to what the model sees at inference time through the Qwen3-Coder Jinja2 chat template. ## Files - `per_turn_plan_qwen3coder.jsonl` — 25,033 per-turn plan samples ## Why This Format ### The Problem Qwen3-Coder uses a **native XML tool call format** that is different from the older Qwen 2.5 JSON-in-XML format: ``` Qwen3-Coder native (correct): Qwen 2.5 (wrong): <tool_call> <tool_call> <function=Bash> {"name": "Bash", "arguments": {...}} <parameter=command> </tool_call> ls /testbed </parameter> </function> </tool_call> ``` LLaMA-Factory's built-in tool handling (`FunctionFormatter` + `QwenToolUtils`) produces the Qwen 2.5 format. Using it would train the model on a tool call format it never sees at inference time. Additionally, LLaMA-Factory's custom template approach (`StringFormatter` for `format_function`) crashes because the `_encode()` method always passes a `thought_words` tuple kwarg that `StringFormatter.apply()` can't handle (it expects only string kwargs). ### The Solution Use **only `human`/`gpt` roles** with the stock **`qwen3_nothink`** template: | Original Role | Converted Role | Content Handling | |-----------------|---------------|-----------------| | `human` | `human` | Unchanged | | `gpt` | `gpt` | Unchanged | | `function_call` | `gpt` | JSON tool calls converted to native XML | | `observation` | `human` | Wrapped in `<tool_response>` tags | This works because: 1. **`format_user` and `format_assistant`** in `qwen3_nothink` are both `StringFormatter` — they do simple `{{content}}` substitution and pass content through unchanged. 2. **No `format_function` is ever called** — there are no `function_call` roles in the data, so `thought_words` never causes a crash. 3. **Pre-formatted content** — tool call XML and `<tool_response>` wrapping are baked into the content strings, which pass through `StringFormatter` unchanged. 4. **Loss masking is preserved** — LLaMA-Factory masks even-index turns (human) and trains on odd-index turns (gpt), same as it would with observation/function_call. ### How It Matches Inference At **inference time** (vLLM + Qwen3-Coder Jinja2 template): ``` <|im_start|>system {system prompt + tools XML} <|im_end|> <|im_start|>user {task description} <|im_end|> <|im_start|>assistant ← model generates tool calls I'll explore the codebase. <tool_call> <function=Bash> <parameter=command> ls /testbed </parameter> </function> </tool_call> <|im_end|> <|im_start|>user ← vLLM wraps tool result <tool_response> file1.py file2.py </tool_response> <|im_end|> <|im_start|>assistant ← model generates response Here's what I found... <|im_end|> ``` At **training time** (LLaMA-Factory + `qwen3_nothink` template): ``` <|im_start|>system {system prompt + tools XML} ← tools baked into system (same XML) <|im_end|> <|im_start|>user {task description} ← human turn (masked, not trained) <|im_end|> <|im_start|>assistant I'll explore the codebase. ← gpt turn (TRAINED) <tool_call> <function=Bash> <parameter=command> ls /testbed </parameter> </function> </tool_call> <|im_end|> <|im_start|>user <tool_response> ← human turn (masked, not trained) file1.py file2.py </tool_response> <|im_end|> <|im_start|>assistant Here's what I found... ← gpt turn (TRAINED) <|im_end|> ``` The bytes between `<|im_start|>assistant\n` and `<|im_end|>` are **identical** in both paths. Verified with `validate_final.py` (8,212 turns, 0 mismatches) and `compare_training_vs_inference.py` (uses LLaMA-Factory's actual Python classes to render and compare). ## Conversion Generated from the original ShareGPT dataset using: ```bash python scripts/convert_to_qwen3coder_format.py \ --input sft_dataset_distill_02/training/per_turn_plan.jsonl \ --output sft_dataset_distill_02/training/v2_human_gpt/per_turn_plan_qwen3coder.jsonl \ --roles human-gpt \ --stats ``` Key transformations: - `function_call` JSON content → Qwen3-Coder native XML, role → `gpt` - `observation` content → wrapped in `<tool_response>\n{content}\n</tool_response>\n`, role → `human` - `tools` JSON field → rendered via actual Jinja2 template, baked into system prompt, field removed ## LLaMA-Factory Configuration ```yaml dataset: ccr_sft_plan_qwen3coder_v2 template: qwen3_nothink ``` Registered in `LlamaFactory/data/dataset_info.json` as `ccr_sft_plan_qwen3coder_v2`. ## Validation ```bash # 1. Validate training format matches inference format (no GPU needed) python scripts/validate_final.py \ --input sft_dataset_distill_02/training/v2_human_gpt/per_turn_plan_qwen3coder.jsonl \ --num-samples 50 # 2. Compare using LLaMA-Factory's actual template classes (no GPU needed) /scratch_mfs/kishan/LlamaFactory/.venv/bin/python3 \ scripts/compare_training_vs_inference.py \ --input sft_dataset_distill_02/training/v2_human_gpt/per_turn_plan_qwen3coder.jsonl # 3. CPU dry run through LLaMA-Factory (uses Qwen3-1.7B, no GPU) cd /scratch_mfs/kishan/LlamaFactory llamafactory-cli train examples/train_balanced/qwen3coder_ccr_sft_plan_dryrun_cpu.yaml ```

# v2_human_gpt — 适用于Qwen3-Coder的监督微调(Supervised Fine-Tuning,SFT)训练数据 ## 概述 本目录包含可直接适配**原生LLaMA-Factory**(无需修改)的SFT训练数据,且生成的**Token**与模型在推理阶段通过Qwen3-Coder的Jinja2对话模板所接收的Token完全一致。 ## 文件 - `per_turn_plan_qwen3coder.jsonl` — 共25033条每轮规划样本 ## 为何采用该格式 ### 问题所在 Qwen3-Coder采用**原生XML工具调用格式**,与旧版Qwen 2.5的JSON内嵌XML格式存在差异: Qwen3-Coder原生格式(正确): Qwen 2.5格式(错误): <tool_call> <tool_call> <function=Bash> {"name": "Bash", "arguments": {...}} <parameter=command> </tool_call> ls /testbed </parameter> </function> </tool_call> LLaMA-Factory内置的工具处理模块(`FunctionFormatter` + `QwenToolUtils`)会生成Qwen 2.5格式的数据,若使用该格式训练,模型将在推理阶段从未见过的工具调用格式下进行训练。 此外,LLaMA-Factory的自定义模板方案(针对`format_function`的`StringFormatter`)会出现崩溃,原因是`_encode()`方法始终会传递一个`thought_words`元组关键字参数,而`StringFormatter.apply()`无法处理该参数(其仅支持字符串类型的关键字参数)。 ### 解决方案 仅使用**`human`/`gpt`角色**,并采用原生**`qwen3_nothink`**模板,对应关系如下: | 原始角色 | 转换后角色 | 内容处理方式 | |----------------|------------|----------------------------------| | `human` | `human` | 保持不变 | | `gpt` | `gpt` | 保持不变 | | `function_call`| `gpt` | 将JSON格式的工具调用转换为原生XML | | `observation` | `human` | 用`<tool_response>`标签包裹内容 | 该方案可行的原因如下: 1. **`qwen3_nothink`模板中的`format_user`与`format_assistant`均为`StringFormatter`**:仅执行简单的`{{content}}`替换,且直接透传原始内容。 2. **无需调用`format_function`**:数据中不存在`function_call`角色,因此不会触发`thought_words`参数导致的崩溃问题。 3. **内容预格式化**:工具调用XML与`<tool_response>`标签包裹的内容已预先嵌入到内容字符串中,可通过`StringFormatter`直接透传。 4. **损失掩码保持一致**:LLaMA-Factory会对偶数索引轮次(human角色)进行掩码,仅在奇数索引轮次(gpt角色)上进行训练,这与使用observation/function_call角色时的行为完全一致。 ### 与推理流程的一致性 在**推理阶段**(vLLM + Qwen3-Coder Jinja2模板): <|im_start|>system {system prompt + tools XML} <|im_end|> <|im_start|>user {task description} <|im_end|> <|im_start|>assistant ← 模型生成工具调用 I'll explore the codebase. <tool_call> <function=Bash> <parameter=command> ls /testbed </parameter> </function> </tool_call> <|im_end|> <|im_start|>user ← vLLM封装工具返回结果 <tool_response> file1.py file2.py </tool_response> <|im_end|> <|im_start|>assistant ← 模型生成回复内容 Here's what I found... <|im_end|> 在**训练阶段**(LLaMA-Factory + `qwen3_nothink`模板): <|im_start|>system {system prompt + tools XML} ← 工具配置已嵌入系统提示词(格式一致) <|im_end|> <|im_start|>user {task description} ← human轮次(掩码,不参与训练) <|im_end|> <|im_start|>assistant I'll explore the codebase. ← gpt轮次(参与训练) <tool_call> <function=Bash> <parameter=command> ls /testbed </parameter> </function> </tool_call> <|im_end|> <|im_start|>user <tool_response> ← human轮次(掩码,不参与训练) file1.py file2.py </tool_response> <|im_end|> <|im_start|>assistant Here's what I found... ← gpt轮次(参与训练) <|im_end|> `<|im_start|>assistant `与`<|im_end|>`之间的字节内容在训练与推理流程中**完全一致**。该结论已通过`validate_final.py`(覆盖8212轮次,无任何不匹配项)与`compare_training_vs_inference.py`(使用LLaMA-Factory的实际Python类进行渲染与对比)验证。 ## 数据转换流程 本数据集由原始ShareGPT数据集通过以下命令生成: bash python scripts/convert_to_qwen3coder_format.py --input sft_dataset_distill_02/training/per_turn_plan.jsonl --output sft_dataset_distill_02/training/v2_human_gpt/per_turn_plan_qwen3coder.jsonl --roles human-gpt --stats 关键转换步骤包括: - 将`function_call`的JSON格式内容转换为Qwen3-Coder原生XML格式,并将角色修改为`gpt` - 将`observation`的内容用`<tool_response> {content} </tool_response> `标签包裹,并将角色修改为`human` - 将`tools`字段的JSON配置通过实际的Jinja2模板渲染为XML格式,嵌入系统提示词中,并移除原`tools`字段 ## LLaMA-Factory配置 yaml dataset: ccr_sft_plan_qwen3coder_v2 template: qwen3_nothink 该数据集已在`LlamaFactory/data/dataset_info.json`中注册为`ccr_sft_plan_qwen3coder_v2`。 ## 验证流程 bash # 1. 验证训练格式与推理格式匹配(无需GPU) python scripts/validate_final.py --input sft_dataset_distill_02/training/v2_human_gpt/per_turn_plan_qwen3coder.jsonl --num-samples 50 # 2. 使用LLaMA-Factory的实际模板类进行对比(无需GPU) /scratch_mfs/kishan/LlamaFactory/.venv/bin/python3 scripts/compare_training_vs_inference.py --input sft_dataset_distill_02/training/v2_human_gpt/per_turn_plan_qwen3coder.jsonl # 3. 通过LLaMA-Factory执行CPU离线训练(使用Qwen3-1.7B,无需GPU) cd /scratch_mfs/kishan/LlamaFactory llamafactory-cli train examples/train_balanced/qwen3coder_ccr_sft_plan_dryrun_cpu.yaml

提供机构:
PGCodeLLM
二维码
社区交流群
二维码
科研交流群
商业服务