遇见数据集

Bas95/fineinstructions_nemotron

收藏
Hugging Face2026-04-29 更新2026-05-31 收录
官方服务:

资源简介:

该数据集名为FineInstructions,是一个大规模合成指令-答案对数据集,包含约10亿+(~1B+)合成指令-答案对,或约300B个词元(tokens)。这些数据是通过FineInstructions流水线生成的,该流水线基于Nemotron-CC预训练语料库(来自CommonCrawl的高质量文档子集)的原始预训练文档运行。数据以.parquet文件格式存储在data文件夹中,每个文件都有一个对应的judge-*.json文件,包含基于Likert评分(1-5,5为最高质量)的自动质量评分。

This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline. The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). Each .parquet file in the data folder has a corresponding judge-*.json file that contains an automatic judgement score of the quality of the synthetic instruction-answer pair on a Likert score (1-5) where 5 is the highest-quality.

提供机构:
Bas95
二维码
社区交流群
二维码
科研交流群
商业服务