官方服务:
资源简介:
Documentation of the Hocank project (DoBeS project)
应用场景:
相关数据集
IGC-2022-1
冰岛语Gigaword语料库(IGC)是一个包含近24亿词的8个子语料库集合。其中部分子语料库以开放许可(CC-BY)发布,包括期刊、法律、新闻1、议会、社交和维基等子集。每个子集可能包含两个或多个子语料库。此数据集以jsonl格式存储,每个文件包含一篇新闻文章或议会会议等。每行是一个JSON对象,包含文档全文、随机生成的ID、作者、转换时间戳、原始XML文件ID、发布时间戳、标题、段落和句子偏移
Hugging Face2025-05-28 更新90
Komnzo text corpus
This dataset contains a zip-file of the most up to date version of the Komnzo text corpus. The original footage files can be found here at Zenodo or under: Döhler, C. 2010-2015. DoBeS Documentat
Zenodo2024-09-02 更新50
SINM-OTP5 - East Kola'a Ridge, Honiara
unknown unknown The Solomon Islands National Museum licenses the use of this recording for non-commercial educational purposes. Please acknowledge them and PARADISEC in any use of this material. Digit
Research Data Australia40
commoncrawl/gneissweb-annotation-host-testing-v1
GneissWeb注释数据集是一个应用于Common Crawl语料库的质量和类别注释数据集,由IBM Research的GneissWeb方法提供支持。该数据集支持对医疗、教育、技术和科学领域网络内容的精确过滤,便于为研究项目、语言模型和专业应用构建高质量语料库。数据集包含两个层次的注释粒度:主机级别(整个域的聚合统计)和URL级别(单个URL分类)。数据集利用了IBM公开提供的GneissWe
Hugging Face2025-12-11 更新50
CH3-091700Xi - 1700 items, recorded by X
1700 items, recorded by X. Language as given: Mro Khimi [cmr]
Research Data Australia50



