遇见数据集

nebius/gpt-oss-120b-Infinity-Instruct-0625

收藏
Hugging Face2026-03-02 更新2026-04-05 收录
官方服务:

资源简介:

--- license: cc-by-4.0 task_categories: - text-generation language: - en configs: - config_name: default data_files: - split: train path: data/train-* dataset_info: features: - name: conversation list: - name: content dtype: string - name: role dtype: string - name: generated_message struct: - name: annotations dtype: 'null' - name: audio dtype: 'null' - name: content dtype: string - name: function_call dtype: 'null' - name: reasoning_content dtype: string - name: refusal dtype: 'null' - name: role dtype: string - name: tool_calls sequence: 'null' - name: finish_reason dtype: string splits: - name: train num_bytes: 6061060422 num_examples: 659358 download_size: 3572636672 dataset_size: 6061060422 --- # gpt-oss-120b-Infinity-Instruct-0625 ## Dataset Description This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside gpt-oss-120b as the target model. The dataset was created by generating responses to the prompts from [Infinity-Instruct-0625](https://huggingface.co/datasets/BAAI/Infinity-Instruct) with [openai/gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b) at temperature=1. For more details on the training methodology and results, see our paper: [LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding](https://arxiv.org/abs/2602.23881). ## Dataset Structure - **Format**: parquet - **Rows**: 659,358 ## Usage ```python from datasets import load_dataset dataset = load_dataset("nebius/gpt-oss-120b-Infinity-Instruct-0625") ``` ## License The dataset is released under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) ## Citation ``` @misc{samarin2026lklosses, title = {LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding}, author = {Alexander Samarin and Sergei Krutikov and Anton Shevtsov and Sergei Skvortsov and Filipp Fisin and Alexander Golubev}, year = {2026}, eprint = {2602.23881}, archivePrefix = {arXiv}, primaryClass = {cs.LG}, url = {https://arxiv.org/abs/2602.23881} } ```

许可证:CC BY 4.0 任务类别: - 文本生成(text-generation) 语言: - 英语(en) 配置项: - 配置名称:default 数据文件: - 拆分集:train 路径:data/train-* 数据集信息: 特征: - 名称:conversation 列表类型: - 名称:content 数据类型:字符串 - 名称:role 数据类型:字符串 - 名称:generated_message 结构体类型: - 名称:annotations 数据类型:空值 - 名称:audio 数据类型:空值 - 名称:content 数据类型:字符串 - 名称:function_call 数据类型:空值 - 名称:reasoning_content 数据类型:字符串 - 名称:refusal 数据类型:空值 - 名称:role 数据类型:字符串 - 名称:tool_calls 序列类型:空值 - 名称:finish_reason 数据类型:字符串 拆分集: - 名称:train 字节大小:6061060422 样本数量:659358 下载大小:3572636672 数据集总大小:6061060422 # gpt-oss-120b-Infinity-Instruct-0625 ## 数据集说明 本数据集属于用于推测式解码(speculative decoding)研究的LK-Speculators数据集集合,包含66万条提示词-响应对,专为训练草稿模型而设计,该草稿模型将与作为目标模型的gpt-oss-120b配合使用。本数据集通过使用openai/gpt-oss-120b(温度参数设为1)对来自Infinity-Instruct-0625的提示词生成响应而构建。如需了解训练方法与实验结果的更多细节,请参阅我们的论文:《LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding》(https://arxiv.org/abs/2602.23881)。 ## 数据集结构 - **数据格式**:Parquet - **样本总数**:659358 ## 使用方法 python from datasets import load_dataset dataset = load_dataset("nebius/gpt-oss-120b-Infinity-Instruct-0625") ## 许可证 本数据集采用知识共享署名4.0(CC BY 4.0)许可证发布,详情请访问:https://creativecommons.org/licenses/by/4.0/ ## 引用格式 @misc{samarin2026lklosses, title = {LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding}, author = {Alexander Samarin and Sergei Krutikov and Anton Shevtsov and Sergei Skvortsov and Filipp Fisin and Alexander Golubev}, year = {2026}, eprint = {2602.23881}, archivePrefix = {arXiv}, primaryClass = {cs.LG}, url = {https://arxiv.org/abs/2602.23881} }

提供机构:
nebius
二维码
社区交流群
二维码
科研交流群
商业服务