遇见数据集

qwen3.7-max-split-formatted

收藏
魔搭社区2026-08-18 更新2026-07-15 收录
官方服务:

资源简介:

# Qwen Agent Thinking Online Distillation Rows This dataset contains cumulative assistant-turn training rows prepared for online logit distillation of Qwen-style agent models, plus a small set of no-tools chat rows to reduce tool-call overbias. Each row is a rendered-chat-ready conversation prefix ending at a target assistant turn. The trainer uses all prior messages as context and applies loss only to the final assistant span. ## Dataset Details - Source trace repo: `armand0e/qwen3.7-max-pi-traces` - Chat stabilizer repo: `TeichAI/claude-4.5-opus-high-reasoning-250x` - Format: JSON Lines - Split: `train` - Rows: `1,485` - Agent/tool rows: `1,235` - Chat-only no-tools rows: `250` - Source sessions: `47` - Maximum assistant step: `108` - Language: English - Intended use: online teacher/student logit distillation for agentic coding traces ## Files - `train.jsonl`: cumulative distillation rows The dataset card YAML explicitly maps `train.jsonl` to the `train` split using the `configs` block. ## Row Schema Each JSONL row contains: - `messages`: list of chat messages ending at the target assistant message - `tools`: available-tool schema for agent rows; omitted for chat-only stabilizer rows - `metadata`: provenance and slicing metadata Important metadata fields: - `source`: source dataset repository - `session_id`: source trace session id - `assistant_steps`: total assistant turns in the source session - `step_index`: cumulative assistant-turn index for this row - `target_message_index`: index of the final assistant target in `messages` - `prefix_message_count`: number of messages included in the row - `stage`: `online_logit_distill` for agent trace rows or `chat_stabilizer` for no-tools chat rows - `tools_scope`: records whether the row carries the corpus-global tool schema or omits tools ## Construction Agent rows were built from Pi-style agent trace JSONL sessions. For each session, the converter emits cumulative rows ordered by assistant step: 1. all session prefixes ending at assistant turn 1 2. all session prefixes ending at assistant turn 2 3. all session prefixes ending at assistant turn 3 4. continuing until the longest session is exhausted The ordering is curriculum-style across the corpus, not grouped by full session. Assistant `reasoning_content`, text content, tool calls, and tool results are preserved when present. Agent rows receive the same available-tool schema: `bash`, `read`, `write`, and `edit`. The chat stabilizer rows come from `TeichAI/claude-4.5-opus-high-reasoning-250x`. They omit `tools` entirely, so the model also sees normal no-tool conversations during distillation. ## Loading ```python from datasets import load_dataset dataset = load_dataset("armand0e/qwen3.7-max-split-formatted", "default", split="train") print(dataset[0]) ``` ## Intended Use This dataset is intended for online logit distillation where a teacher model is run live during training and the student is trained with a mix of: - cross-entropy on the recorded target assistant tokens - KL divergence against the teacher distribution on those same target positions Non-target context tokens should be masked out during loss computation. ## Limitations - The dataset is generated from trace data and may contain failed tool calls, partial attempts, or behavior specific to the source agent environment. - The rows are not independently human-verified. - Long cumulative rows may exceed a trainer's maximum sequence length; trainers should filter out rows that do not fully fit rather than cutting away prior context. - License and redistribution constraints may depend on the source trace collection and any underlying generated content. Treat this dataset as a derived training artifact rather than a clean public benchmark. ## Citation No formal citation is available. If used, cite the source dataset repository and this derived dataset repository.

提供机构:
maas
创建时间:
2026-06-10
二维码
社区交流群
二维码
科研交流群
商业服务