Annoy-PyEdu-Rs
收藏资源简介:
该数据集是Annoy项目资源集合的一个子集,名为Annoy-PythonEdu-Rs。由于合作者的合规要求,仅发布了此PythonEdu-Rs子集。数据集采用完全基于大语言模型(LLM)的方法合成,旨在生成可靠且富有表现力的响应,以克服传统基于模板方法在表达性和泛化性上的限制。数据集的许可证为ODC-BY。用户还可以访问经过处理的原始数据版本。
This dataset is part of the resource collection of the 'Annoy' project, specifically the 'Annoy-PythonEdu-Rs' subset. Due to the compliance requirements of our collaborators, only this PythonEdu-Rs subset has been released. This dataset is synthesized entirely using large language model (LLM)-based methods, aiming to generate reliable and expressive responses to overcome the limitations of traditional template-based approaches in terms of expressiveness and generalizability. The dataset is licensed under ODC-BY. Users can also access the processed original data version.
数据集概述
数据集名称:Annoy-PythonEdu-Rs(Annoy-PyEdu-Rs)
发布页面:https://huggingface.co/datasets/liu12123456/Annoy-PyEdu-Rs
关联资源:
- 原始处理后数据:https://huggingface.co/datasets/liu12123456/Annoy-PyEdu-Rs-Raw
- 所属资源集合:https://huggingface.co/collections/liu12123456/specx-67a978e28fd926b56a4f55a2
相关论文与项目:
- 论文:https://huggingface.co/papers/xxxx.xxxxx
- 项目页面:https://specx.github.io/
- 代码仓库:https://github.com/swis-laiguni/Annoy
数据集背景与构建方法
该数据集是 Annoy 项目资源的一部分,旨在生成可靠的执行轨迹作为响应。由于完全可执行代码在输入预测和自动构建轨迹方面存在局限(例如难以获得确定性反向函数、轨迹受限于预设计模板),研究者采用 基于LLM的全合成方法,使用 DeepSeek-V2.5 生成所有期望的响应。DeepSeek-V2.5 具有顶级性能且成本极低。
重要说明:因合作方的合规要求,目前仅公开发布 PythonEdu-Rs 子集(即本页面数据集),而非完整数据集。
数据集许可
- 许可协议:ODC-BY
相关模型
本数据集配套训练了多个模型,分为 Annoy 和 Annoy++ 两个系列,每个系列包含 Stage 1 和 Stage 2 两个训练阶段。基座模型包括:




