Arlow-Constellations
收藏资源简介:
该数据集是一个用于预训练Arlow的预训练集,由多个流式数据源混合并打乱而成。数据集包含41个训练分片,每个分片包含500万条样本,每条样本包含两个字段:'text'(文本内容)和'source'(数据来源)。总数据量约为1.27 TB,下载大小约为430.67 GB。数据集适用于文本生成任务,采用odc-by许可协议。数据来源包括HuggingFaceFW/fineweb、multilingual-mi-llm/pile、openbmb/UltraData-Math、CohereLabs/aya_dataset等多个公开数据集。
This dataset is a pre-training dataset for Arlow, which is compiled from multiple streaming data sources and shuffled. It contains 41 training shards, with each shard holding 5 million samples. Each sample includes two fields: 'text' (text content) and 'source' (data source). The total data volume is approximately 1.27 TB, and the download size is about 430.67 GB. This dataset is suitable for text generation tasks and is released under the ODC-By license. Its data sources cover multiple public datasets such as HuggingFaceFW/fineweb, multilingual-mi-llm/pile, openbmb/UltraData-Math, CohereLabs/aya_dataset, and more.
Arlow-Constellations 数据集概述
数据集基本信息
- 数据集名称: Arlow-Constellations
- 托管地址: https://huggingface.co/datasets/yuchenxie/Arlow-Constellations
- 许可证: odc-by
- 主要任务类别: 文本生成
数据配置与结构
- 默认配置名称: default
- 数据特征:
text: 字符串类型source: 字符串类型
- 数据划分: 数据集包含41个训练子集(train_0 至 train_40),每个子集均为独立的训练分割。
数据规模
- 总下载大小: 430,667,875,271 字节
- 总数据集大小: 1,268,148,030,465 字节
- 总样本数量: 205,000,000 条(每个训练子集包含5,000,000条样本,共41个子集)
数据来源与构成
该数据集是一个经过混洗的混合数据集,由以下流式数据源混合而成:
- HuggingFaceFW/fineweb (sample-350BT)
- multilingual-mi-llm/pile
- openbmb/UltraData-Math (UltraData-Math-L1)
- CohereLabs/aya_dataset inputs
- CohereLabs/aya_dataset targets
- HuggingFaceFW/fineweb-edu (sample-350BT)
- PleIAs/common_corpus
- openbmb/UltraData-Math (UltraData-Math-L3-Multi-Style-Synthetic)
- bigcode/the-stack
- nvidia/Nemotron-CC-Math-v1 (4plus_MIND)
数据格式说明
- 每一行数据包含且仅包含两个列:
text和source。 - 该数据集用于预训练 Arlow 模型。




