Turkish-Python-Expert-Dataset-PY_CORE-DATA_STRUCTURES-130k
收藏资源简介:
该数据集是一个面向土耳其语使用者的 Python 专家指令数据集,包含约 130,000 条高质量指令对,用于自然语言代码生成和调试。V2 版本在原有 60,000 条基础 Python 指令基础上,新增了 70,000 条高级/专家级场景,涵盖网络架构、Python 内部机制及混合技能问题。数据集按渐进式软件工程课程设计,分为 27 个类别,包括基础语法、数据结构、算法、面向对象、函数式编程、文件 I/O、数据库、Web 框架、网络、并发、异步、测试、安全、DevOps、Shell、性能优化、调试、通信、系统工具、集成项目,以及 V2 新增的 Web 抓取、任务队列、高级类型系统和 Python 内部机制。数据格式为 ChatML/Alpaca 指令格式,语言为纯正土耳其语,无翻译痕迹。适用于训练土耳其语代码助手、高级系统机器人以及学术研究中的 Code-LLM 基准测试。数据集由 Hakan Ttkr (Bysismo) 设计,采用 MIT 许可证。
This dataset is a Python expert instruction dataset for Turkish speakers, containing approximately 130,000 high-quality instruction pairs for natural language code generation and debugging. The V2 version adds 70,000 advanced/expert-level scenarios on top of the original 60,000 basic Python instructions, covering network architecture, Python internals, and hybrid skill problems. The dataset is designed according to a progressive software engineering curriculum and is divided into 27 categories, including basic syntax, data structures, algorithms, object-oriented programming, functional programming, file I/O, databases, web frameworks, networking, concurrency, asynchronous programming, testing, security, DevOps, shell, performance optimization, debugging, communication, system tools, integrated projects, and V2 additions such as web scraping, task queues, advanced type systems, and Python internals. The data format is ChatML/Alpaca instruction format, and the language is pure Turkish without translation traces. It is suitable for training Turkish code assistants, advanced system bots, and Code-LLM benchmarks in academic research. The dataset is designed by Hakan Ttkr (Bysismo) and licensed under MIT.
数据集概述
Turkish Python Expert Dataset - 130K (V2) 是一个面向土耳其语软件开发者和人工智能助手的高质量Python指令微调数据集,包含约130,000条训练数据,专注于自然语言下的Python代码生成与调试任务。
基本信息
- 语言: 土耳其语 (tr)
- 任务类型: 文本生成、问答
- 标签: python、coding、networking、instruction-tuning、synthetic-data
- 许可协议: MIT License(开放学术和商业使用)
- 数据集格式: ChatML / Alpaca Instruction Format
- 开发者: Hakan Ttkr (Bysismo)
版本与规模
- V1(基础版): 包含60,000条“基础Python”问题
- V2(更新版): 新增70,000条高级/专家级场景,涵盖网络架构、Python核心机制及“混合”技能场景
- 当前总行数: 约130,000+ 行
核心内容与分类
本次发布聚焦于 PY_CORE(Python核心:基础语法、内置函数)和 DATA_STRUCTURES(数据结构:列表、字典、集合、deque等),包含14,000个独立问题。数据集设计的课程体系共涵盖27个类别,包括:
- 基础层: PY_CORE、DATA_STRUCTURES、ALGORITHMS(进行中)
- 进阶层: OOP、FUNCTIONAL、FILE_IO、DATABASE、WEB_BASICS
- 框架层: FASTAPI、FLASK、DJANGO、NETWORKING、CONCURRENCY、ASYNC
- 工程层: TESTING、SECURITY、DEVOPS、SHELL、PERFORMANCE、DEBUGGING
- 专家层(V2新增): WEB_SCRAPING、TASK_QUEUES、ADVANCED_TYPING、PYTHON_INTERNALS
新增特性
- 混合交叉引擎: V2引入了跨多类别组合的复杂架构问题(例如FastAPI + Celery + Pytest + eBPF组合场景)
- 自然土耳其语: 所有内容以流畅、非翻译腔的土耳其语编写,模型可直接给出清晰答案
应用场景
- 训练土耳其语代码助手(类似GitHub Copilot)的监督微调(SFT)阶段
- 构建高级系统级AI模型,处理网络命名空间、eBPF、内存管理等CPython底层问题
- 作为土耳其语Code-LLM性能基准测试的基础数据资源




