ravwojdyla/datakit-tier2-skewed-v2
收藏资源简介:
Datakit Tier2 Skewed Synthetic是一个合成的、重尾文档数据集,专为压力测试Marin datakit管道(包括规范化、最小哈希、模糊去重、整合和标记化等步骤)而生成,针对高达256 MB的文档长度异常值进行设计。这是一个用于持续集成/管道压力测试的数据集,而非训练语料库,旨在覆盖长尾代码路径,这些路径在FineWeb-Edu的冒烟测试中未涉及。数据集基于FineWeb-Edu的10BT样本分割生成,采用三模式混合分布:70%的对数正态分布(平均约5 KB)、30%的帕累托分布(alpha=1.1, scale=2 KB),以及100个在128-256 MB范围内随机注入的大型文档。文件格式为Parquet,包含id(字符串,16位十六进制字符)和content(字符串,UTF-8文本正文)两列。许可证为ODC-By 1.0。
A synthetic, heavy-tailed-document dataset generated for stress-testing the Marin datakit pipeline (normalize / minhash / fuzzy_dups / consolidate / tokenize) against doc-length outliers up to 256 MB. This is a CI / pipeline-stress dataset, not a training corpus. It exists to exercise long-tail code paths that the FineWeb-Edu smoke ferry doesnt cover. The dataset is generated from the FineWeb-Edu 10BT sample split using a 3-mode mixture distribution: 70% log-normal (mean ~5 KB), 30% Pareto alpha=1.1 scale=2 KB, plus 100 mega docs uniformly in [128, 256] MB injected at random positions. The schema includes Parquet files with id (string, 16 hex chars) and content (string, UTF-8 text body) columns. Licensed under ODC-By 1.0.



