遇见数据集

cds-jb/synthweb-qwen3-8b

收藏
Hugging Face2026-05-12 更新2026-05-31 收录
官方服务:

资源简介:

该数据集名为qwen3-8b-fineweb-rollouts-100k,是一个基于Qwen3-8B模型生成的FineWeb前缀延续文本集合。它包含来自HuggingFaceFW/fineweb的sample-10BT子集的100,000个源文档,每个文档前缀最多生成15个延续文本。数据集经过模式崩溃过滤和截断处理,以移除模型从延续FineWeb散文模式切换到后训练脚手架模式(如链式思维脚手架、理解问答或摘要结尾)的部分。处理流程包括:基于长度分布的文档池筛选、模型生成、模式崩溃过滤(如重复n-gram检测)、后训练标记截断(在检测到特定模式时截断generated_text)和长度下限过滤(保留至少400字符的延续)。数据集模式与源parquet相同,但增加了tail_truncated_at_marker列,用于记录触发截断的正则模式。截断而非丢弃整个延续文本的目的是保留有用的散文部分,避免因后续漂移而损失数据。

The dataset named qwen3-8b-fineweb-rollouts-100k is a collection of Qwen3-8B continuations of FineWeb prefixes. It includes 100,000 source documents from the sample-10BT subset of HuggingFaceFW/fineweb, with up to 15 rollouts per prefix. The dataset is filtered for mode-collapse and truncated at detached-end markers to remove parts where Qwen3-8B switches from continuing FineWeb prose into post-training scaffold modes (e.g., chain-of-thought scaffold, comprehension Q&A, or summary coda). The processing pipeline involves: document pool filtering to match length distribution, model generation, mode-collapse filtering (e.g., n-gram repetition detection), truncation at post-training markers (cutting generated_text at detected patterns), and a length floor (dropping rollouts with fewer than 400 characters). The schema matches the source parquet with an added tail_truncated_at_marker column to record the regex pattern that triggered truncation. Truncation is preferred over dropping entire rollouts to preserve useful prose portions before drift occurs.

提供机构:
cds-jb
二维码
社区交流群
二维码
科研交流群
商业服务