KantaHayashiAI/ClimbMix-Ja-Initial64-Training-Data
收藏资源简介:
ClimbMix-Ja Initial64 350M Artifacts是一个公共备份数据集,用于存储初始的64个ClimbMix-Ja候选运行的数据和配置。该数据集基于nvidia/nemotron-climb-proxy-models 350M基础模型,转换为Megatron-LM TE兼容的检查点,并使用KantaHayashiAI/ClimbLab-Ja训练语料库,该语料库被聚类为cluster_01到cluster_20。训练序列长度为1024,每个候选运行进行6500次迭代,全局批次大小为304,每个候选训练token数为2,023,424,000,所有候选总训练token数为129,499,136,000。数据集内容包括聚类的JSONL文件、Megatron-LM索引数据集文件、清单文件以及混合物定义和候选脚本。匹配的检查点存储库为KantaHayashiAI/ClimbMix-Ja-350M-Initial64-Checkpoints。
ClimbMix-Ja Initial64 350M Artifacts is a public backup dataset for storing the data and configurations of the initial 64 ClimbMix-Ja candidate runs. It is based on the nvidia/nemotron-climb-proxy-models 350M base model, converted to a Megatron-LM TE-compatible checkpoint, and uses the KantaHayashiAI/ClimbLab-Ja training corpus, which is clustered into cluster_01 to cluster_20. The training sequence length is 1024, with 6500 iterations per candidate, a global batch size of 304, 2,023,424,000 tokens per candidate, and a total of 129,499,136,000 tokens across all candidates. The dataset contents include clustered JSONL files, Megatron-LM indexed dataset files, manifest files, and mixture definitions with candidate scripts. The matching checkpoint repository is KantaHayashiAI/ClimbMix-Ja-350M-Initial64-Checkpoints.




