JoeyLLM/canada-dataset-1b
收藏资源简介:
这是一个从Common Crawl中提取的更大规模加拿大网络文本语料库的10亿token代表性样本。该样本作为JoeyLLM项目区域英语语言模型研究的一部分发布,覆盖了2013年至2025年的Common Crawl转储数据,包含约1,463,015个文档(行),压缩大小为2.89 GB。数据集旨在用于语言模型预训练、加拿大英语使用(如拼写、地名、机构)的领域适应实验、Common Crawl衍生语言模型数据集研究,以及JoeyLLM数据管道输出的检查和可重复性。数据集中每个文档都经过清理,包括文本、ID、URL、日期等字段,并通过分层随机抽样方法从完整语料库(2100亿token,未发布)中选取,以保持代表性。
This is a 1-billion-token representative sample of a larger-scale Canadian web text corpus extracted from Common Crawl. Released as part of regional English language model research for the JoeyLLM project, this sample covers Common Crawl dump data spanning 2013 to 2025, containing approximately 1,463,015 documents (lines) with a compressed size of 2.89 GB. This dataset is intended for language model pre-training, domain adaptation experiments focused on Canadian English usage (such as spelling conventions, toponyms, and institutional terminology), research on Common Crawl-derived language model datasets, as well as inspection and reproducibility of outputs from the JoeyLLM data pipeline. Each document in the dataset has been cleaned, including fields such as text, ID, URL, and date, and was selected from the unreleased full corpus (210 billion tokens) via stratified random sampling to maintain representativeness.



