遇见数据集

josephmayo/public-curated-coding-data

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

这是一个名为公共收集编码数据的数据集,它混合了来自多个公共来源的编码相关数据,包括GitHub源代码、Reddit编程讨论、Hacker News讨论、Stack Exchange问答以及从上游Hugging Face数据集镜像的Claude Opus推理代码数据。数据集已规范化为提示/响应对格式(每行包含prompt和response字符串字段),专门用于语言模型实验。数据经过去重和代码相关性过滤处理,总计包含2,703行数据。数据集支持英语,主要任务类别为文本生成,标签涵盖编码、公共数据、GitHub、StackExchange、Hackernews、Reddit和推理等领域。请注意,数据来源多样,许可协议不统一,需参考相关文件了解具体条款。

This dataset is named Public Collected Coding Data. It consists of mixed-origin public coding data normalized into prompt/response pairs for language-model experiments. The data sources include GitHub source code, Reddit programming discussions, Hacker News discussions, Stack Exchange Q&A, and Claude Opus reasoning code data mirrored from an upstream Hugging Face dataset. The dataset is formatted as JSONL rows, each containing exactly prompt and response string fields. It has undergone deduplication and code-relevance filtering, with a total of 2,703 rows across all configs. The language is English, the primary task category is text-generation, and tags include coding, public-data, github, stackexchange, hackernews, reddit, and reasoning. Note that the data originates from multiple sources with varying licenses; refer to accompanying documentation for specific terms.

提供机构:
josephmayo
二维码
社区交流群
二维码
科研交流群
商业服务