相关数据集
Corpus of the Contemporary Lithuanian Language
This corpus includes Lithuanian texts (mostly newspapers but also fiction, non-fiction, and specialised magazines) published between 1990 and 2008. The corpus is encoded in TEI. Non-linguistic metadat
SSH Open MarketPlace2023-10-13 更新190
Estonian Treebank annotated with coreference relations
This corpus contains newspaper texts plus one scientific medical text. The corpus is available for download from META-SHARE (CELR distribution).
SSH Open MarketPlace2023-11-14 更新170
salmankhanpm/Corpous_Telugu_Tokenizer
这是一个包含大量文本字符串的数据集,具体为一个名为context的大型字符串特征。数据集被划分为训练集,共有7729706个示例,大小为40659788654字节。数据集的下载大小为15630491429字节。
Hugging Face2025-10-18 更新90
International Corpus of English (Australia contribution is ICE-AUS)
The ICE-AUS corpus is a 1 m.word corpus of transcribed spoken and written Australian English from 1992-1995. Its internal structure with 500 samples (60% speech, 40%writing) matches that of other ICE
Research Data Australia180
The corpus of older Slovenian narrative prose PriLit 1.0
This corpus contains texts of older Slovenian narrative prose by 12 authors. The corpus is available for download from the CLARIN.SI repository.
SSH Open MarketPlace2024-11-13 更新110



