官方服务:
资源简介:
Language Acquisition corpus
应用场景:
相关数据集
IndicCorp
IndicCorp 是一个大型单语语料库,拥有大约 90 亿个代币,涵盖 12 种主要的印度语言。它是通过在几个月的时间内发现和抓取数千个网络资源(主要是新闻、杂志和书籍)而开发的。涵盖的语言:阿萨姆语、孟加拉语、英语、古吉拉特语、印地语、卡纳达语、马拉雅拉姆语、马拉地语、奥里亚语、旁遮普语、泰米尔语、泰卢固语语料库格式:语料库是一个大型文本文件,每行包含一个句子。公开发布的版本被随机洗牌、去标记
OpenDataLab2026-07-12 更新330
Segakorpus: Doktoritööd Corpus of Estonian scientific texts
Korpus sisaldab 5 miljonit sõna eestikeelset teaduskirjandust: doktoritööd (2,3 miljonit sõna) ja teadusartiklid. TEI P5 XML märgendus, UTF8 kodeering. More info at...
B2FIND60
dgskorpus_fra_16
This is one out of 165 recordings done in the DGS-Korpus project (http://dgs-korpus.de) from 2010 to 2012 to collect a corpus of DGS (German Sign Language). The recording took...
B2FIND40
ITEM 5 -- Class corpora per project
* One zip file per application domain* application domains contain the zip's for each project
DataCite Commons2020-08-27 更新60
motazsaad/comparableWikiCoprus: v1.0
Full Changelog: https://github.com/motazsaad/comparableWikiCoprus/commits/v1.0
NIAID Data Ecosystem60



