遇见数据集

Nemotron-Pretraining-Specialized-v1.2

收藏
魔搭社区2026-06-14 更新2026-07-15 收录
官方服务:

资源简介:

# Nemotron-Pretraining-Specialized-v1.2 ## Dataset Description: The [Nemotron-Pretraining-Specialized-v1.2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.2) dataset is part of the [Nemotron Pretraining Data](https://huggingface.co/collections/nvidia/nemotron-pre-training-datasets) collection of pretraining datasets. Designed for the [NVIDIA Nemotron 3](https://huggingface.co/collections/nvidia/nvidia-nemotron-v3) family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions. Note: These are new datasets, not replacements. They are meant to be used together with the previously released datasets [Nemotron-Pretraining-Specialized-v1.1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.1) and [Nemotron-Pretraining-Specialized-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1). This dataset is ready for commercial use. ## Dataset Details: For more details, please see the [NVIDIA Nemotron 3 Ultra tech report](https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf). This dataset has the following subsets: - **Nemotron-Pretraining-Fact-Seeking**: Fact-seeking questions generated from [Finewiki](https://huggingface.co/datasets/HuggingFaceFW/finewiki). In an ablation on an intermediate Nemotron 3 Nano base model checkpoint, training with this data improved a proxy SimpleQA score from 40.2 to 50.2. - **Nemotron-Pretraining-Moral-Scenarios**: In the [SFT data](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1) we previously released, we included multiple-choice questions about moral scenarios. These questions were constructed using situations and norms from Moral Stories and actions from Social Chemistry. We sampled a subset of these examples and created a chain-of-thought version using Qwen3-235B-A22B-Thinking-2507. - **Nemotron-Pretraining-Generative**: Diverse large-scale synthetic questions with generative answers. - **Nemotron-Pretraining-Multiple-Choice**: Diverse large-scale synthetic questions and answers in multiple choice format. For more details about how the **Nemotron-Pretraining-Generative** and **Nemotron-Pretraining-Multiple-Choice** subsets were created, see [Task-Seeded Synthetic Q&A Generation for Nemotron Pretraining](https://huggingface.co/blog/nvidia/task-seeded-sdg). The table below shows the number of tokens and the model used to generate these subsets: | Subset | Tokens (M) | Models | License | | --- | --- | --- | --- | | **Nemotron-Pretraining-Fact-Seeking** | 35029.6 | Qwen3-30B-A3B-Instruct-2507 | cc-by-4.0 | | **Nemotron-Pretraining-Moral-Scenarios** | 15.2 | Mixtral-8x22B-v0.1, Qwen3-235B-A22B-Thinking-2507 | cc-by-4.0, additional information: MIT | | **Nemotron-Pretraining-Generative** | 685.2 | DeepSeek-v3, Qwen3-235B-A22B | cc-by-4.0 or cc-by-2.0: see 'license' entry for each sample | | **Nemotron-Pretraining-Multiple-Choice** | 6099.6 | DeepSeek-v3, Qwen3-235B-A22B | cc-by-4.0 or cc-by-2.0: see 'license' entry for each sample | The columns are as follows: - **text**: The **primary data field,** containing the content to be used for pretraining. - **license**: The license(s) governing the sample (e.g., ‘cc-by-4.0’). - **metadata**: A dictionary detailing the following: - **category**: Data type ('Nemotron-Pretraining-Generative' or 'Nemotron-Pretraining-Multiple-Choice'). - **models_used**: Models used to generate the data (e.g., ''). - **uuid**: The unique identifier for this dataset entry. ## Dataset Owner(s): NVIDIA Corporation ## Dataset Creation Date: 05/18/2026 ## Version: [Nemotron-Pretraining-Specialized-v1.2](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.2) Previous Version(s): - [Nemotron-Pretraining-Specialized-v1.1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.1) - [Nemotron-Pretraining-Specialized-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1) - [Nemotron-Pretraining-SFT-v1](https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1) Relationship to Previous Version(s): This is an **extension** of the previously released datasets, containing new data. They are meant to be used together. ## License/Terms of Use: The **Nemotron-Pretraining-Fact-Seeking** dataset is licensed under the [Creative Commons Attribution 4.0 International License](https://creativecommons.org/licenses/by/4.0/) (CC-BY-4.0). The **Nemotron-Pretraining-Moral-Scenarios** dataset is licensed under the [Creative Commons Attribution 4.0 International License](https://creativecommons.org/licenses/by/4.0/) (CC-BY-4.0). Additional Information: [MIT License](https://mit-license.org/). For the **Nemotron-Pretraining-Multiple-Choice** and **Nemotron-Pretraining-Generative** datasets, please see the 'license' entry for each sample. The majority of the datasets is licensed under the [Creative Commons Attribution 4.0 International License](https://creativecommons.org/licenses/by/4.0/) (CC-BY-4.0). Some samples are licensed under the [Creative Common Attribution 2.0 Generic License](https://creativecommons.org/licenses/by/2.0/legalcode.en) (CC-BY-2.0). Each user is responsible for checking the content of datasets and the applicable licenses and determining if suitable for the intended use. This dataset contains synthetic data created using the following models: DeepSeek-v3, Mixtral-8x22B-v0.1, Qwen3-30B-A3B-Instruct-2507, Qwen3-235B-A22B, Qwen3-235B-A22B-Thinking-2507. If the **Nemotron-Pretraining-Multiple-Choice** or **Nemotron-Pretraining-Generative** datasets are used to create, train, fine-tune, or otherwise improve an AI model, which is distributed or made available, such AI model may be subject to redistribution and use requirements in the [DeepSeek License Agreement](https://huggingface.co/deepseek-ai/DeepSeek-V3/blob/main/LICENSE-MODEL). ## Intended Usage: The Nemotron-Pre-Training-Specialized-v1.2 Dataset is intended to be used by the community to continue to improve open models. ## Dataset Characterization **Data Collection Method** - Synthetic: Synthetic generation using large language models (DeepSeek-v3, Mixtral-8x22B-v0.1, Qwen3-30B-A3B-Instruct-2507, Qwen3-235B-A22B, Qwen3-235B-A22B-Thinking-2507). **Labeling Method** - Not Applicable ## Dataset Format Modality: Text Format: Parquet ## Dataset Quantification Record Count: 599.5M samples Measurement of Total Data Storage: 53.6 GB ## Reference(s): If you use our dataset in your research, please cite our [NVIDIA Nemotron 3 Ultra tech report](https://research.nvidia.com/labs/nemotron/files/NVIDIA-Nemotron-3-Ultra-Technical-Report.pdf). For more details on **Nemotron-Pretraining-Generative** and **Nemotron-Pretraining-Multiple-Choice**, please see [Task-Seeded Synthetic Q&A Generation for Nemotron Pretraining](https://huggingface.co/blog/nvidia/task-seeded-sdg). ## Ethical Considerations: NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse. Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).

提供机构:
maas
创建时间:
2026-06-06
二维码
社区交流群
二维码
科研交流群
商业服务