官方服务:
资源简介:
Sample of diachronic corpus
历时语料库(diachronic corpus)样本
应用场景:
相关数据集
gsarti/magpie
MAGPIE语料库(Haagsma et al. 2020)是一个大规模的意义标注语料库,包含潜在的习语表达(PIEs),基于英国国家语料库(BNC)。潜在的习语表达类似于习语表达,但也包括习语表达的字面用法,例如“我在一天结束时下班”中的“一天结束时”。该数据集版本反映了Dankers等人(2022)在研究中使用的过滤子集,用于研究NMT模型如何表示PIEs。作者使用了37k个标注为完全比喻或字
Hugging Face2022-10-27 更新450
DD1-1029 - Words from Texts 2
Elicitation about the lexical meaning of various words in the corpus, mainly from Jacob's transcripts. The recorder's maximum filesize was exceeded so it split the session into two recordings.. Langua
Research Data Australia70
MCL - Multifunctional Computational Lexicon of Contemporary Portuguese
MCL is a 26,443 lemma Frequency Lexicon with 140,315 tokens, with the minimum lemma frequency of 6, extracted from CORLEX, a contemporary Portuguese corpus (16,210,438 words). CORLEX is a subcorpus of
DataCite Commons2022-06-01 更新70
Connecting Conditionals (Reuneker 2022; dissertation)
Scripts (Python, R) and data (corpus data from CGN and SoNaR) belonging to the PhD dissertation 'Connecting Conditionals: A corpus-based approach to conditional constructions in Dutch' by Alex Reuneke
DataCite Commons2025-07-03 更新130
CCC1-Rembarrnga - Rembarrnga_Corpus
corpus files for ANNIS corpus viewer Note that there is a set of files for each corpus. The main basic text is in the file that includes 'text' in the title.. Language as given:
Research Data Australia70



