遇见数据集

farida5gaber/stackoverflow-posts

收藏
Hugging Face2026-04-26 更新2026-05-03 收录
官方服务:

资源简介:

该数据集包含了2023年6月14日之前提交到StackOverflow(一个编程问答社区)的所有帖子,并以Markdown文本格式进行存储。数据集规模庞大,包含约6000万个帖子,总数据量约为35GB,文本字符数达到约650亿。数据来源于Internet Archive的StackExchange数据转储。每个记录对应一个特定类型的帖子(如问题、答案、标签维基等),原始数据转储的顺序由于处理脚本的并行性并未完全保留。帖子内容的主要字段Body以Markdown格式存储,替代了原始的HTML格式,便于文本处理和分析。数据集还包含其他元数据字段,如ID、帖子类型、分数、查看次数、标题、内容许可证、创建日期、标签等,适用于问答、文本生成和文本到文本生成等自然语言处理任务。

This dataset contains all posts submitted to StackOverflow before the 14th of June 2023 formatted as Markdown text. The dataset contains ~60 Million posts, totaling ~35GB in size and ~65 billion characters of text. The data is sourced from Internet Archive StackExchange Data Dump. Each record corresponds to one post of a particular type. Original ordering from the data dump is not exactly preserved due to parallelism in the script used to process the data dump. The markdown content of each post is contained in the Body field. The dataset includes fields such as Id, PostTypeId, AcceptedAnswerId, ParentId, Score, ViewCount, Body, Title, ContentLicense, FavoriteCount, CreationDate, LastActivityDate, LastEditDate, LastEditorUserId, OwnerUserId, and Tags, making it suitable for tasks like question-answering, text-generation, and text2text-generation.

提供机构:
farida5gaber
二维码
社区交流群
二维码
科研交流群
商业服务