danish-foundation-models/faroese-dynaword
收藏资源简介:
法罗语Dynaword是一个收集自多个领域的法罗语自由文本数据集的集合。所有数据集均采用开放许可证,并被认为适用于训练大型语言模型。数据集包含58.89K个样本,总计6.50M个令牌(使用Llama 3分词器),平均文档长度为110.43个令牌(范围从2到28.77K)。数据来源包括维基百科(百科全书领域,5.45M令牌)、Ravnursson ASR语料库的标准化法罗语转录(朗读领域,794.44K令牌)以及POS标注的Sosialurin报纸文本(新闻领域,263.51K令牌)。数据集持续开发,会随着新数据的可用性而更新,并鼓励社区贡献。
The Faroese dynaword is a collection of Faroese free-form text datasets from various domains. All of the datasets in the Faroese Dynaword are openly licensed and deemed permissible for training large language models. The dataset contains 58.89K samples with 6.50M tokens (using the Llama 3 tokenizer), and an average document length of 110.43 tokens (ranging from 2 to 28.77K). Sources include Wikipedia (encyclopedic domain, 5.45M tokens), normalized Faroese transcripts from the Ravnursson ASR corpus (readaloud domain, 794.44K tokens), and POS-tagged Sosialurin newspaper text (news domain, 263.51K tokens). It is continually developed, with updates as new datasets become available, and encourages community contributions.




