遇见数据集

JoeyLLM/new-zealand-dataset-1b

收藏
Hugging Face2026-05-13 更新2026-05-31 收录
官方服务:

资源简介:

新西兰网络文本—1B词元样本是一个更大规模已清理新西兰网络文本语料库的代表性样本,源自Common Crawl。该样本包含约10亿个词元,覆盖2013年至2025年,来自109个Common Crawl转储,包含1,749,277行文档,压缩大小为2.97 GB。数据集用于支持区域英语语言模型研究,特别是新西兰英语的预训练和领域适应,包括拼写、毛利语借词、地名、机构等区域语言特征。数据经过语言过滤、质量过滤、国家归属和去重等处理,但可能仍包含噪声和个人信息。

A 1-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM projects ongoing research into regional English language models. The sample represents approximately 1.65% of the full internal New Zealand corpus by token count, while preserving coverage across the Common Crawl dumps and years represented in the full cleaned corpus.

提供机构:
JoeyLLM
二维码
社区交流群
二维码
科研交流群
商业服务