遇见数据集

JoeyLLM/australian-dataset-1b

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

这是一个来自Common Crawl的更大规模清洗过的澳大利亚网络文本语料库的10亿token代表性样本。该样本与JoeyLLM项目正在进行的区域英语语言模型研究一起发布。

A 1-billion-token representative sample of a much larger cleaned Australian web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM projects ongoing research into regional English language models.

提供机构:
JoeyLLM
二维码
社区交流群
二维码
科研交流群
商业服务