遇见数据集

XuehangCang/e_style_code

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

一个中文代码数据集,目标是为“易语言风格”的Python代码生成与指令跟随任务提供训练样本。数据集中每条样本都包含一个通用代码需求(question)、对应的简要思路(thinking),以及一段可直接运行或稍作调整即可运行的中文命名Python实现(answer)。数据集包含613条训练样本,并配套有风格规范文档,用于约束中文命名、注释、函数设计和格式,以统一代码风格。设计目标包括训练生成中文命名Python代码的指令模型、构建“易语言风格”的代码补全和转换能力,以及研究中文编程表达和风格迁移。使用方式可通过Hugging Face Datasets加载,适用于基础代码生成、风格迁移和中文编程表达任务,但建议结合更通用的数据集和代码校验措施。

A Chinese code dataset designed to provide training samples for Easy Language style Python code generation and instruction-following tasks. Each sample in the dataset includes a general code requirement (question), a corresponding brief thinking process (thinking), and a directly runnable or slightly adjustable Chinese-named Python implementation (answer). The dataset contains 613 training samples and comes with a companion style specification document to standardize Chinese naming, comments, function design, and formatting. The design goals include training instruction models for generating Chinese-named Python code, building Easy Language style code completion and conversion capabilities, and researching Chinese programming expression and style transfer. It can be loaded via Hugging Face Datasets and is suitable for basic code generation, style transfer, and Chinese programming expression tasks, but it is recommended to combine with more general datasets and code validation measures.

提供机构:
XuehangCang
二维码
社区交流群
二维码
科研交流群
商业服务