遇见数据集

abhinav0231/Sarvam-105b-Distill-100k

收藏
Hugging Face2026-05-21 更新2026-04-12 收录
官方服务:

资源简介:

Sarvam 105B Distill 100K v5是一个从Sarvam 105B大语言模型蒸馏得到的推理数据集,包含约10万条英文单轮对话样本。数据集采用四种不同模式组织:thinking(原生推理保留模式,包含思考过程和回答)、sharegpt(兼容ShareGPT的对话格式)、chatml(预格式化ChatML文本,可直接用于分词器流水线)和simple_qa(扁平化模式,适用于监督微调和分析)。数据涵盖编程与计算机科学、数学、科学与STEM、逻辑与形式推理、语言与写作、历史地理与公民学、经济学与金融、法律与伦理、健康与医学、创意规划与开放式问题等10个领域,难度分为简单、中等和困难三个级别。数据集包含训练集(92040条)、验证集(1917条)和测试集(1918条)三个分割。

Sarvam 105B Distill 100K v5 is a reasoning dataset distilled from the Sarvam 105B large language model, containing approximately 100,000 English single-turn dialogue samples. The dataset is organized in four different schemas: thinking (native reasoning-preserving schema with separate thinking and response fields), sharegpt (ShareGPT-compatible conversation schema), chatml (preformatted ChatML text for direct tokenizer pipelines), and simple_qa (flat schema for supervised finetuning and analytics). The data covers 10 domains including coding and computer science, mathematics, science and STEM, logic and formal reasoning, language and writing, history and geography and civics, economics and finance, law and ethics, health and medicine, and creative planning and open-ended questions. The difficulty levels are categorized as easy, medium, and hard. The dataset includes three splits: train (92,040 samples), validation (1,917 samples), and test (1,918 samples).

提供机构:
abhinav0231
二维码
社区交流群
二维码
科研交流群
商业服务