遇见数据集

musabc/nanogpt-tr-v5-data

收藏
Hugging Face2026-05-16 更新2026-05-31 收录
官方服务:

资源简介:

nanogpt-tr-v5数据集是一个用于土耳其语语言模型预训练的分词数据集,专门为训练200M参数的土耳其语语言模型(V5版本)设计。数据集包含多个层级的数据文件:v5_stage1.bin(Web层,来自OSCAR、mC4、论坛和FineWeb-HQ等来源,约2.94B词元)、v5_stage2.bin(Medium层,来自BellaTurca、Cosmos、CulturaX、Havadis和Cosmopedia等来源,约9.03B词元)、v5_stage3.bin(Premium层,来自Wiki、Wikisource、Tezler、Akademik、FinePDFs和Özenli等来源,约2.97B词元),以及v5_val.bin(验证集,从所有三个阶段的最后1%数据中提取,约150M词元)。此外,数据集还包括一个BPE分词器文件(tokenizer-tr-v5.json),词汇量为32K,基于Stage3数据训练。数据以uint16词元ID格式存储,可通过Numpy memmap高效读取,适用于大规模自然语言处理任务。

The nanogpt-tr-v5 dataset is a tokenized dataset for Turkish language model pre-training, specifically designed for training a 200M-parameter Turkish language model (V5 version). The dataset includes multiple tiered data files: v5_stage1.bin (Web tier, from sources such as OSCAR, mC4, forums, and FineWeb-HQ, approximately 2.94B tokens), v5_stage2.bin (Medium tier, from sources including BellaTurca, Cosmos, CulturaX, Havadis, and Cosmopedia, approximately 9.03B tokens), v5_stage3.bin (Premium tier, from sources like Wiki, Wikisource, Tezler, Akademik, FinePDFs, and Özenli, approximately 2.97B tokens), and v5_val.bin (validation set, extracted from the last 1% of all three stages, approximately 150M tokens). Additionally, the dataset includes a BPE tokenizer file (tokenizer-tr-v5.json) with a vocabulary size of 32K, trained on Stage3 data. The data is stored in uint16 token ID format and can be efficiently read via Numpy memmap, making it suitable for large-scale natural language processing tasks.

提供机构:
musabc
二维码
社区交流群
二维码
科研交流群
商业服务