遇见数据集

Voidreaper2026/coding-master-dataset

收藏
Hugging Face2026-05-06 更新2026-05-31 收录
官方服务:

资源简介:

Coding Master Dataset 是一个大规模编码指令调优数据集,采用 ShareGPT 对话格式,通过汇编多个开源数据源并经过去重处理构建而成。该数据集包含 766,987 条记录,以 JSONL 格式存储,遵循 Apache 2.0 许可证。它专为文本生成和问答任务设计,适用于代码相关的指令调优和微调,覆盖 Python、JavaScript、Go 等编程语言,规模在 10 万到 100 万条之间。数据来源包括 CodeX-2M-Thinking、python-code-dataset-500k、StackPulse 高质量子集、CodeFeedback-Filtered-Instruction、secure_programming_dpo 和 Go 编码 JSONL 文件,确保了多样性和覆盖面。数据集结构包含 conversations(ShareGPT 格式的人机对话轮次)、source(来源数据集)、source_id(唯一标识符)和 metadata_json(附加元数据)等字段,便于使用和扩展。

A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated. It contains 766,987 records in JSONL format, licensed under Apache 2.0, and is designed for text-generation and question-answering tasks, suitable for code-related instruction tuning and fine-tuning across programming languages like Python, JavaScript, and Go, with a size category of 100K<n<1M. Sources include CodeX-2M-Thinking, python-code-dataset-500k, StackPulse high-quality subset, CodeFeedback-Filtered-Instruction, secure_programming_dpo, and Go coding JSONL files, ensuring diversity and coverage. The schema features fields such as conversations (ShareGPT format human/gpt turns), source (origin dataset), source_id (unique record identifier), and metadata_json (additional metadata) for ease of use and extensibility.

提供机构:
Voidreaper2026
二维码
社区交流群
二维码
科研交流群
商业服务