azmx-6t-data
收藏资源简介:
AZMX One 6.6T 预训练语料库是一个大规模、高质量、经过清洗、去重和过滤的预训练数据集,专为 AZMX One(一个拥有 1.87B 参数的开放印度语言大模型)而构建。该数据集规模约为 6.6 万亿 token,以 uint32 v2-tokenizer 分片格式存储。数据集的构建部分基于 Modal 平台。该语料库主要用于训练多语言印度语言模型,适用于自然语言处理领域的预训练任务。
The AZMX One 6.6T pre-training corpus is a large-scale, high-quality, cleaned, deduplicated and filtered pre-training dataset built specifically for AZMX One, an open-source Indian language large language model (LLM) with 1.87 billion parameters. This dataset has a scale of approximately 6.6 trillion tokens, and is stored in the sharded format using the uint32 v2-tokenizer. Part of the development of this dataset was based on the Modal platform. This corpus is primarily intended for training multilingual Indian language models, and is applicable to pre-training tasks in the field of natural language processing.
数据集概述
AZMX One 6.6T 预训练语料库 是一个面向印地语(Indic)大型语言模型训练的预处理数据集,专为 AZMX One(一个 1.87B 参数的开源 Indic LLM)设计。
核心特性
- 规模:6.6T tokens 的预训练数据
- 处理流程:已完成清洗(clean)、去重(deduped)和过滤(filtered)
- 构建工具:部分基于 Modal 平台构建
- 数据格式:uint32 格式的 v2-tokenizer 分片(shards)
适用场景
该数据集主要用于 AZMX One 模型的预训练阶段,面向印地语(Indic)语言处理任务,可作为开源 Indic 大语言模型的训练语料来源。




