wizzense/Usenet-Corpus-1980-2013
收藏资源简介:
Usenet Corpus 1980–2013是一个大型、经过精心整理的Usenet存档数据集,专为AI训练设计。它包含4.08亿条经过清洗和去重的Usenet帖子,覆盖1980年至2013年期间的18,347个新闻组,总计1031亿个令牌。数据集源自最大的私有Usenet语料库之一,并经过严格处理,适用于现代AI训练用例。关键统计包括:总帖子数408,236,288,总令牌数1031亿,英语内容占比96.6%,主要新闻组层次结构包括ALT(57%)、REC(16%)、COMP(10%)、SOC(8%)和SCI(3.2%)。每个记录包含文本、新闻组名称、日期、主题、作者(已匿名处理)和ID等字段。数据集经过广泛清洗,移除了alt.binaries.*和成人内容、二进制编码,并进行了去重和敏感内容过滤,确保数据质量。预期用途包括LLM的预训练和持续预训练、领域适应(技术、科学和历史话语)、对话和长上下文建模、语言学研究以及去偏和时间语言研究。数据集不适用于重新识别、分析或垃圾邮件生成。访问方式包括学术/非商业用途的免费预览和商业用途的许可选项。
Usenet Corpus 1980–2013 is one of the largest curated Usenet archives available for AI training, containing 408 million cleaned and deduplicated Usenet posts spanning 1980–2013 across 18,347 newsgroups, with a total of 103.1 billion tokens. It is sourced from one of the largest privately held Usenet corpora and has been rigorously processed for modern AI training use cases. Key statistics include total posts of 408,236,288, total tokens of 103.1 billion, English content at 96.6%, and major hierarchies such as ALT (57%), REC (16%), COMP (10%), SOC (8%), and SCI (3.2%). Each record includes fields like text, group, date, subject, author (with PII redacted), and ID. The corpus has undergone extensive cleaning, including removal of alt.binaries.* and adult content, stripping binaries, deduplication, sensitive content filtering, and PII redaction. Intended uses include pre-training and continued pre-training of LLMs, domain adaptation for technical, scientific, and historical discourse, dialogue and long-context modeling, linguistic research, and debiasing and temporal language studies. It is not intended for re-identification, profiling, or spam generation. Access is available for academic/non-commercial use via a free preview on Hugging Face, with commercial licensing options for full corpus access.




