Corpus of Slovene linguistic scientific writing JezKor
收藏资源简介:
This corpus contains a collection of linguistic scientific writing in the Slovenian language. It consists of 43 monographs published between 2009 and 2022 by Fran Ramovš institute of Slovenian language and Založba ZRC, 267 papers published in the journal "Jezikoslovni zapiski" and 28 papers published in the journal "Slovenski jezik". Note that the texts were obtained directly from PDFs, so they contain various types of noise. The corpus is linguistically annotated with the CLASSLA pipeline (https://github.com/clarinsi/classla) on the levels lemmatisation, MULTEXT-East Version 6 morphosyntactic descriptions, Universal Dependencies part-of-spech and morphological features, and named entities. It is distributed in CoNLL-U and vertical file format, one file for each text. Text metadata consists of the author(s), title and year of publication. The corpus is available for download from the CLARIN.SI repository as well as for online browsing through the noSketch Engine and KonText concordancers.
本语料库收录斯洛文尼亚语语言学学术写作文本。其包含由斯洛文尼亚语言弗朗·拉莫夫什研究所与ZRC出版社(Založba ZRC)于2009年至2022年间出版的43部专著、发表于《Jezikoslovni zapiski》期刊的267篇论文,以及发表于《Slovenski jezik》期刊的28篇论文。需注意,文本均直接从PDF文件提取,因此包含多种类型的噪声。 本语料库已通过CLASSLA流水线(CLASSLA pipeline,https://github.com/clarinsi/classla)完成多维度语言学标注,覆盖词形还原、MULTEXT-East第6版形态句法描述、通用依存句法(Universal Dependencies)词性与形态特征,以及命名实体标注。语料库以CoNLL-U与垂直文件格式分发,每份文本对应一个独立文件。文本元数据包含作者、标题与出版年份信息。 本语料库可从CLARIN.SI知识库下载,同时支持通过noSketch Engine与KonText上下文检索工具进行在线浏览。



