Annoy-PyEdu-Rs
收藏资源简介:
Annoy-PythonEdu-Rs 是一个基于代码执行轨迹和自然语言推理合成的教育编程数据集,是完整 Annoy 数据集的 PythonEdu-Rs 子集,专门针对 Python 教育场景设计。该数据集通过对原始数据 Annoy-PythonEdu-Rs-Raw 进行转换和注释处理而生成,原始数据源采用 Apache License 2.0 许可证。数据集中的响应内容完全使用 DeepSeek-V2.5 大型语言模型合成,旨在解决传统基于模板的代码执行轨迹方法在表达力和泛化能力上的局限性。该数据集适用于代码生成、程序理解、教育编程辅助等任务,为研究如何结合可执行代码与自然语言推理提供了实验数据。数据集遵循 Apache-2.0 开源协议。
Annoy-PythonEdu-Rs is an educational programming dataset synthesized based on code execution traces and natural language inference. This dataset is the PythonEdu-Rs subset of the full Annoy dataset, specifically designed for Python education scenarios. It is generated by transforming and annotating the raw source data Annoy-PythonEdu-Rs-Raw, whose original data source is licensed under the Apache License 2.0. The response content within this dataset is fully synthesized using the DeepSeek-V2.5 Large Language Model, aiming to address the limitations of traditional template-based code execution trace methods in terms of expressiveness and generalization ability. This dataset is applicable to tasks such as code generation, program comprehension, educational programming assistance, etc., providing experimental data for research on combining executable code with natural language inference. The dataset is released under the Apache-2.0 open-source license.
数据集概述
- 数据集名称: Annoy-PythonEdu-Rs (Annoy-PyEdu-Rs)
- 许可证: Apache-2.0
- 发布地址: https://huggingface.co/datasets/dfdfdg5667/Annoy-PyEdu-Rs
数据来源与构建
- 原始数据: 该数据集是
Annoy-PythonEdu-Rs-Raw的转换/标注衍生版本,原始数据可在 dfdfdg5667/Annoy-PyEdu-Rs-Raw 获取。 - 构建方法: 采用完全基于大语言模型的方法,使用 DeepSeek-V2.5 合成所有期望的响应。DeepSeek-V2.5 不对模型输出主张权利,因此不增加额外的输出许可证,但使用者仍需遵守 DeepSeek-V2.5 的许可条款。
- 发布范围: 由于合作方的合规要求,当前仅发布完整数据集中的 PythonEdu-Rs 子集。
数据用途与动机
- 虽然拥有完整的可执行代码理论上可以生成可靠的执行轨迹作为响应,但存在两个挑战:
- 输入预测的确定性逆函数难以获得。
- 自动构建的轨迹受限于预设计模板,缺乏自由形式自然语言推理的表达力和泛化能力。
- 因此,采用 DeepSeek-V2.5(性能优越且成本极低)来合成所有期望的响应。
使用许可
- 数据集许可证为 Apache-2.0,需遵循该许可条款,并在适用情况下保留上游/源通知。
- 数据集中的响应由 DeepSeek-V2.5 合成,使用者需遵守 DeepSeek-V2.5 的许可条款。




