VytautoDidziojoUniversitetas/LT_Morphosyntax_Corpus_SIMAS
收藏资源简介:
SIMAS属于通用标注语料库类别,包含2000年至2025年间撰写和出版的文本,代表当代标准立陶宛书面语。语料库由四个部分组成,反映不同语域:小说、学术写作、行政文本和新闻(发布在立陶宛国家广播电视门户LRT.lt上的文本)。语料库经过自动形态和句法标注,并由语言学家审核。自动形态标注使用工具“Morfuoklis”,自动句法分析使用适用于立陶宛语的国际工具UDPipe。标注遵循国际通用依赖(UD)标准,形态分析还参考了Jablonskis标注标准和MULTEXT-East,句法分析应用了立陶宛语语法标注指南。语料库以单个文件分发,格式为CoNLL-U,每行代表一个单词(标记),用制表符分隔的列提供其词元、词性、形态特征和与其他词的句法关系的详细信息。
SIMAS belongs to the category of general annotated corpora of written language. The corpus consists of texts written and published between 2000 and 2025 and represents original contemporary standard written Lithuanian. The corpus is composed of four sections reflecting different registers: fiction, academic writing, administrative texts, and journalism (texts published on the Lithuanian National Radio and Television portal LRT.lt). The corpus was automatically annotated morphologically and syntactically and subsequently reviewed by linguists. Automatic morphological annotation was carried out using the tool “Morfuoklis”. Automatic syntactic parsing was performed using the international tool UDPipe adapted for the Lithuanian language. The corpus is annotated according to the international Universal Dependencies (UD) standard. The following standards were also used for morphological analysis: Jablonskis morphological annotation standard and MULTEXT-East. For syntactic analysis, the Universal Dependencies Standard: the Lithuanian Syntactic Annotation Guidelines was applied. The corpus is distributed as a single file. The file is provided in CoNLL-U format. CoNLL-U is a text-based data format used to annotate natural language texts according to the UD standard. Each line in a file represents a single word (token), with tab-separated columns providing detailed information about its lemma, part of speech, morphological features, and syntactic relations to other words in the sentence.



