JoeyLLM/uk-dataset-1b
收藏资源简介:
这是一个来自更大规模清理后的英国网络文本语料库的10亿标记代表性样本,该语料库源自Common Crawl。此样本与JoeyLLM项目关于区域英语语言模型的持续研究一同发布。样本覆盖了2013年至2025年的Common Crawl转储数据,保留了完整语料库中语言和年份的覆盖范围,约占完整内部英国语料库标记数的0.14%。数据集用于语言模型的预训练、英国英语用法的领域适应实验,以及Common Crawl衍生语言模型数据集的研究。
This is a representative 1-billion-token sample derived from a larger, cleaned UK web text corpus sourced from Common Crawl. This sample is released alongside the ongoing regional English language model research of the JoeyLLM project. The sample covers Common Crawl dumps from 2013 to 2025, retains the linguistic and temporal coverage of the full corpus, and accounts for approximately 0.14% of the total token count of the full internal UK corpus. This dataset is intended for pretraining of language models, domain adaptation experiments for British English usage, and research on Common Crawl-derived language model datasets.



