遇见数据集

raj2708/wikipedia-en

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

这是一个经过预处理的英文维基百科数据转储版本,基于enwiki-latest,适用于大型语言模型预训练。数据集包含约570万篇文章,总大小为38.2 GB,语言为英语,采用CC-BY-SA 4.0许可证。数据以JSON格式存储,每行代表一篇文章,包含唯一ID、来源、类型、语言、完整文章纯文本以及元数据(如URL、标题、许可证、爬取时间、质量评分、词元计数和单词计数)。该版本经过清理和预处理,便于直接用于机器学习任务。

A clean, preprocessed version of the English Wikipedia dump (enwiki-latest), ready for LLM pretraining. The dataset contains approximately 5.7 million articles, with a total size of 38.2 GB, in English language, under CC-BY-SA 4.0 license. Data is stored in JSON format, with each row representing an article, including unique ID, source, type, language, full article plain text, and metadata (such as URL, title, license, scraped time, quality score, token count, and word count). This version is cleaned and preprocessed for easy use in machine learning tasks.

提供机构:
raj2708
二维码
社区交流群
二维码
科研交流群
商业服务