Annoy-PyEdu-Rs
收藏资源简介:
Annoy-PythonEdu-Rs是一个专门用于代码生成或编程教育任务的数据集,属于更大规模Annoy数据集的子集。由于合作者的合规要求,目前仅公开了PythonEdu-Rs部分。该数据集采用完全基于大语言模型(LLM)的方法合成所需响应,具体使用DeepSeek-V2.5模型进行生成,这种方法在保证高性能的同时具有较低成本。数据集采用odc-by许可证,并提供了经过处理的原始数据版本(Annoy-PyEdu-Rs-Raw)。
Annoy-PythonEdu-Rs is a dataset designed for code generation or programming education-related tasks, and it is a subset of the larger Annoy dataset. Due to compliance requirements from collaborators, only the PythonEdu-Rs portion is currently publicly released. The dataset employs a fully large language model (LLM)-based approach to synthesize the required responses, specifically using the DeepSeek-V2.5 model for generation, which ensures high performance while maintaining low cost. It is licensed under odc-by, and a processed raw data version (Annoy-PyEdu-Rs-Raw) is also provided.
数据集概述
数据集名称: Annoy-PyEdu-Rs
所属项目: Annoy(论文标题)
发布平台: Hugging Face
许可协议: odc-by(开放数据共享署名许可)
数据集描述:
- 该数据集是完整数据集的一个子集,仅包含 PythonEdu-Rs 部分。
- 数据集的构建采用基于大语言模型(LLM)的合成方法,使用 DeepSeek-V2.5 生成所有期望的响应,原因在于该模型性能顶尖且成本极低。
- 研究动机:拥有完整可执行代码理论上可以生成可靠的执行轨迹作为响应,但存在两个挑战:
- 获得用于输入预测的确定性逆函数不切实际。
- 自动构建的轨迹受限于预设计模板,缺乏自由形式自然语言推理的表达力和泛化能力。
相关资源:
- 原始数据(处理后): safaf4455/Annoy-PyEdu-Rs-Raw
- 项目页面: https://specx.github.io/
- 论文链接: https://huggingface.co/papers/xxxx.xxxxx
- 已发布资源集合: https://huggingface.co/collections/safaf4455/specx-67a978e28fd926b56a4f55a2
- 代码仓库: https://github.com/phoanttheijale/Annoy
相关模型:




