Nemotron-Pretraining-Code-v3
收藏资源简介:
Nemotron-Pretraining-Code-v3数据集是Nemotron预训练数据集合的一部分,旨在提升大语言模型(LLMs)的编码能力。该数据集包含对应原始源代码更新的元数据,是Nemotron-Pretraining-Code-v2和v1源代码语料库的增量扩展,新增了来自GitHub的1.46亿个新文件(约1730亿个词元),数据截止日期为2025年9月30日。该数据集仅包含增量新文件,必须与v2和v1语料库联合使用。数据集采用单一配置Nemotron-Code-Metadata,包含以下列:commit_id(Git提交哈希的前七个字符)、rel_path(相对于仓库根目录的文件路径)和language(检测到的编程语言)。数据形式为文本,存储格式为Parquet,记录数量为1.463亿个样本,总存储量为8.2 GB。数据收集和标注方法均为自动化。数据集适用于文本生成和预训练任务,标签包括文本、预训练、人类、法律和Nemotron_3_Ultra,语言为代码,规模类别在1亿到10亿之间。数据集基于CC-BY-4.0许可证发布,可用于商业用途,由NVIDIA Corporation于2026年5月18日创建。使用该数据集时,建议引用NVIDIA Nemotron 3 Ultra技术报告,并注意伦理考虑,确保符合相关行业和使用场景的要求。
The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron pretraining data collection, designed to enhance the coding capabilities of large language models (LLMs). It includes metadata corresponding to updates in original source code, serving as an incremental expansion of the Nemotron-Pretraining-Code-v2 and v1 source code corpora, adding 146 million new files (approximately 173 billion tokens) from GitHub, with data up to September 30, 2025. This dataset contains only incremental new files and must be used in conjunction with the v2 and v1 corpora. It features a single configuration, Nemotron-Code-Metadata, with columns: commit_id (first seven characters of Git commit hash), rel_path (file path relative to the repository root), and language (detected programming language). The data is in text form, stored in Parquet format, with 146.3 million samples and a total storage size of 8.2 GB. Data collection and annotation methods are automated. The dataset is suitable for text generation and pretraining tasks, with tags including text, pretraining, human, legal, and Nemotron_3_Ultra, language as code, and a scale category between 100 million and 1 billion. It is released under the CC-BY-4.0 license, permissible for commercial use, and created by NVIDIA Corporation on May 18, 2026. When using this dataset, it is recommended to cite the NVIDIA Nemotron 3 Ultra technical report and consider ethical implications to ensure compliance with relevant industry and usage scenario requirements.
数据集概述
- 名称: Nemotron-Pretraining-Code-v3
- 所有者: NVIDIA Corporation
- 创建日期: 2026年5月18日
- 许可证: Creative Commons Attribution 4.0 International (CC-BY-4.0)
- 版本: v3(扩展自 v2 和 v1,需合并使用)
任务与标签
- 任务类别: 文本生成(text-generation)
- 标签: 文本、预训练、人类、合法、Nemotron_3_Ultra
- 语言: 代码(code)
数据规模与格式
- 样本数量: 146.3M(约1.46亿)
- 数据存储大小: 8.2 GB
- 数据格式: Parquet
- 模态: 文本
数据集组成
- 配置名称: Nemotron-Code-Metadata
- 数据拆分: 单拆分(train),包含多个 Parquet 文件(part_00000 至 part_00063)
- 数据列说明:
commit_id: Git 提交哈希的前七位字符rel_path: 相对仓库根目录的文件路径language: 检测到的编程语言
数据来源与构建
- 数据收集方法: 自动化
- 标注方法: 自动化
- 数据来源: GitHub,截止日期为2025年9月30日
- 新增内容: 146M 新文件(173B tokens),仅为增量文件,需与 v2 和 v1 语料库联合使用
用途与引用
- 预期用途: 提升大型语言模型的编码能力,面向社区开源模型改进
- 引用: 使用该数据集时,请引用 NVIDIA Nemotron 3 Ultra 技术报告
伦理与合规
- 商业可用性: 是
- 伦理考量: 建议开发者结合行业和用例检查数据集;质量问题可报告至 NVIDIA 安全漏洞提交页面




