Nemotron-CC
收藏资源简介:
Nemotron-CC是由英伟达创建的一个高质量长时预训练数据集,旨在优化大规模语言模型的训练。该数据集包含6.3万亿个token,其中4.4万亿为全球去重后的原始token,1.9万亿为合成token。数据集的创建过程结合了分类器集成、合成数据重述和减少对启发式过滤器的依赖。Nemotron-CC主要用于解决长token视野训练中的数据质量和数量平衡问题,特别是在训练超过15万亿token的模型时,能够显著提升模型的准确性和多样性。
Nemotron-CC is a high-quality long-duration pre-training dataset developed by NVIDIA to optimize the training of large-scale language models. It contains a total of 6.3 trillion tokens, of which 4.4 trillion are globally deduplicated raw tokens and 1.9 trillion are synthetic tokens. The dataset's creation process integrates classifier ensembles, synthetic data paraphrasing, and reduced reliance on heuristic filters. Nemotron-CC is primarily designed to address the trade-off between data quality and quantity in long-token-context training, and it can significantly improve the accuracy and diversity of models, especially when training models with over 15 trillion tokens.

- 1Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining DatasetNVIDIA · 2024年



