官方服务:
资源简介:
Source: 2 Corinthians (Croatian)
来源:《哥林多后书》(克罗地亚语)
应用场景:
相关数据集
Croatian linguistic training corpus hr500k 2.0
This training corpus contains about 500,000 tokens manually annotated on the levels of tokenisation, sentence segmentation, morphosyntactic tagging, lemmatisation and named entities. About half of the
SSH Open MarketPlace2023-10-13 更新180
The CLASSLA-Stanza model for lemmatisation of standard Croatian 2.1
The model for lemmatisation of standard Croatian was built with the [CLASSLA-Stanza tool](https://github.com/clarinsi/classla) by training on the [hr500k training corpus](http://hdl.handle.net/11356/1
SSH Open MarketPlace2023-10-13 更新130
A Translation of Parts of The Qur'an Into Wolof (NCAC_RDD_TAPE_0006B_SE1)
Teachings of Islam Translation of Qur'an into Wolof Alternative names: Ceesay, Cisse, Seesay, Sissay, Sise, Sisse, Arfang, Foday, Sidibeh, Bakary, Bakari,...
B2FIND40
hrwac
HrWac语料库主要面向克罗地亚语,通过爬取.hr顶级域名构建,包含2011年和2014年的数据。该语料库规模较大,包含数十亿级别的token。数据经过段落级别的去重、变音符号恢复标准化、词性标注和词形还原等处理,并按段落打乱。每段包含URL、域名和语言识别(克罗地亚语vs.塞尔维亚语)等元数据,主要用于文本生成和Masked Language Modeling等任务,并采用CC-BY-SA 3.
OpenCSG2024-07-19 更新140
CHILDES Croatian Kovacevic Corpus
two girls learning Croatian in Zagreb
DataCite Commons2026-04-28 更新60



