遇见数据集

epfl-dlab/llaza-20B

收藏
Hugging Face2026-05-03 更新2026-05-31 收录
官方服务:

资源简介:

Llaza Mixture 20B 是一个用于zip2zip语言模型预训练的20B token子集数据集。它来源于完整的Llaza混合数据集,该数据集在四个顶级领域之间按字节平衡分配:通用领域(占50%,来自HuggingFaceFW/fineweb-edu和sample-100BT)、代码领域(占20%,来自bigcode/the-stack-dedup)、数学领域(占10%,来自HuggingFaceTB/finemath和finemath-3plus)以及多语言领域(占20%,来自epfml/FineWeb2-HQ的20种语言子集)。数据集包含约20,000,000,613个token,18,808,438行,文本字节大小为81.15 GiB,使用Llama 3.1 tokenizer进行token计数。每行数据包含text(文本内容)和source(来源细节)字段。数据集旨在用于语言模型预训练和zip2zip风格训练管道的研究,但继承了上游数据集的质量、过滤、许可和安全性属性,可能包含噪声、重复、敏感或其他不良内容。

This dataset is a 20B-token pretraining subset built for zip2zip language-model pretraining. It is derived from the full Llaza mixture, which is byte-balanced across four top-level domains: General (50% from HuggingFaceFW/fineweb-edu and sample-100BT), Code (20% from bigcode/the-stack-dedup), Math (10% from HuggingFaceTB/finemath and finemath-3plus), and Multilingual (20% from 20 language subsets of epfml/FineWeb2-HQ). The subset contains approximately 20,000,000,613 tokens, 18,808,438 rows, and 81.15 GiB of text bytes, with tokens counted using the Llama 3.1 tokenizer. Each row has a schema with text (content) and source (source detail) fields. It is intended for research on language-model pretraining and zip2zip-style training pipelines, but inherits the quality, filtering, licensing, and safety properties of its upstream datasets, which may include noisy, duplicated, sensitive, or otherwise undesirable content.

提供机构:
epfl-dlab
二维码
社区交流群
二维码
科研交流群
商业服务