Annoy-PyEdu-Rs
收藏资源简介:
Annoy-PythonEdu-Rs 是一个用于代码生成与推理的Python教育相关数据集,是完整Annoy数据集的子集。该数据集采用完全基于大语言模型(LLM)的方法合成,使用DeepSeek-V2.5生成所有期望的响应,以克服传统方法中获取确定性反向函数的不切实际性以及自动构建轨迹受限于预设模板、缺乏自由形式自然语言推理的表达力和泛化能力的问题。由于合作者的合规要求,仅公开发布此PythonEdu-Rs子集。数据集已用于训练多个模型变体,包括基于Qwen 2.5 7B Coder、LLaMA 3.1 8B和DeepSeek v2 Lite Coder等基础模型训练的Annoy和Annoy++模型。数据集许可证为ODC-By。
Annoy-PythonEdu-Rs is a Python education-related dataset for code generation and reasoning, and it is a subset of the complete Annoy dataset. This dataset is synthesized using a fully large language model (LLM)-based approach, with DeepSeek-V2.5 generating all desired responses, to overcome the impracticality of obtaining deterministic inverse functions in traditional methods and the limitations of automated trajectory construction being constrained by preset templates, lacking the expressiveness and generalization ability of free-form natural language reasoning. Due to compliance requirements from collaborators, only this PythonEdu-Rs subset is publicly released. The dataset has been used to train multiple model variants, including Annoy and Annoy++ models trained on base models such as Qwen 2.5 7B Coder, LLaMA 3.1 8B, and DeepSeek v2 Lite Coder. The dataset license is ODC-By.
数据集概述
- 数据集名称:Annoy-PythonEdu-Rs(简称 Annoy-PyEdu-Rs)
- 发布机构:sdadafdaf4546
- 许可证:odc-by(Open Data Commons Attribution License)
数据集来源与背景
- 属于 Annoy 项目资源的一部分,该项目旨在通过基于大语言模型(LLM)的方法合成高质量的执行轨迹响应。
- 使用 DeepSeek-V2.5 作为教师模型自动生成数据,以克服预定义模板的局限性和表达力不足的问题。
- 注意:受合作方合规要求限制,目前仅发布完整数据集中的 PythonEdu-Rs 子集(即本页面)。
数据内容与用途
- 领域:Python 教育(PythonEdu),侧重于有完整可执行代码的教育场景。
- 合成策略:采用完全 LLM 驱动的方式,由 DeepSeek-V2.5 模型生成自由形式的自然语言推理轨迹,而非依赖确定性反向函数或预定义模板。
- 用途:可用于训练或评估模型在代码执行、推理链生成等任务上的能力。
附加资源
- 原始处理数据:可访问 sdadafdaf4546/Annoy-PyEdu-Rs-Raw 获取。
- 相关模型:该数据集配套发布了一系列基于不同基座模型(Qwen 2.5 7B Coder、LLaMA 3.1 8B、DeepSeek v2 Lite Coder)和训练阶段(Stage 1 / Stage 2)的 Annoy 与 Annoy++ 模型,详情可查看 Released Resources。
附注




