官方服务:
资源简介:
Language Acquisition corpus
应用场景:
相关数据集
ELLIPSE-Corpus
ELLIPSE语料库是一个免费提供的语料库,包含约6,500份英语学习者的写作样本,这些样本已被评分用于整体语言熟练度以及与连贯性、句法、词汇、短语学、语法和约定相关的分析熟练度分数。此外,该语料库还提供了语料库中英语学习者的个人和人口统计信息,包括经济状况、性别、年级水平(8-12年级)和种族/民族。该语料库为个人作者提供语言熟练度分数,并旨在推动语料库和NLP方法在评估整体和更精细熟练度特征方
github2024-05-20 更新1610
Wanca 2016, Korp Version
The Korp version of Wanca 2016 is a collection of web corpora in small Uralic languages. The collection is composed of 29 sentence corpora in different languages. The corpora have been collected from
Mendeley Data2024-01-31 更新90
Modal verbs of strong obligation in Scottish Standard English: Corpus data
The dataset contains instances of the (semi-)modal verbs 'must', 'have to', 'need to' and '(have) got to' from nineteen written and spoken genres in the Scottish and British components of the Internat
DataONE2026-01-05 更新70
MultiOpenSubs Corpus
该语料库基于OPUS OpenSubtitles数据,主要用于比较语言学研究。它包含两个子语料库:Multiopensubs_euro(仅包含欧洲语言文本)和Multiopensubs_misc(包含非欧洲语言文本)。每个子语料库包含14种语言,具有统一的翻译单元和令牌数量,以避免数据不一致。
github2023-06-27 更新520
SRBCorp: Corpus of Parliamentary Debates in Serbia
+++++++++++++++++++++++++++++++++++++++++++ The most recent version of this study is available at: https://doi.org/10.5281/zenodo.6521648...
B2FIND70



