TimeML annotated corpus of Estonian newspaper articles
收藏资源简介:
Estonian TimeML Annotated Corpus (ver 2.0) The corpus consists of 80 Estonian newspaper articles (approx. 22,000 word tokens) with manually corrected morphological and dependency syntactic annotations, and with manually added temporal semantic annotations. This corpus is a subcorpus of Estonian Dependency Treebank ( https://github.com/EstSyntax/EDT ). Temporal semantic annotations are based on an adaption of the TimeML specification ( http://www.timeml.org/ ), and consist of EVENT, TIMEX and TLINK annotations. The creation process of the corpus, along with the evaluation of consistency of annotation is described by Orasmaa (2014a, 2014b). Format of the corpus See https://github.com/soras/EstTimeMLCorpus/blob/master/readme.txt for details. Related publications The creation of this corpus and its first version is described in publications: S.Orasmaa (2014a). Towards an Integration of Syntactic and Temporal Annotations in Estonian. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14). S.Orasmaa (2014b). How Availability of Explicit Temporal Cues Affects Manual Temporal Relation Annotation. Human Language Technologies - The Baltic Perspective (215 - 218). IOS Press.
爱沙尼亚语TimeML标注语料库(版本2.0) 本语料库包含80篇爱沙尼亚语报纸文章(约22000个词元(Token)),配有经人工校正的词法形态标注与依存句法标注,同时添加了人工标注的时间语义标注。该语料库为爱沙尼亚语依存树库(Estonian Dependency Treebank,EDT)的子语料库,其来源地址为:https://github.com/EstSyntax/EDT。 时间语义标注基于TimeML规范(http://www.timeml.org/)的适配方案,涵盖事件(EVENT)标注、时间表达式(TIMEX)标注与时间链接(TLINK)标注三类内容。该语料库的构建流程及标注一致性评估方法由Orasmaa(2014a、2014b)详细阐述。 语料库格式 有关该语料库的详细格式说明,请参阅:https://github.com/soras/EstTimeMLCorpus/blob/master/readme.txt。 相关学术出版物 本语料库的构建工作及其首个版本的相关研究发表于以下学术成果: 1. S.Orasmaa(2014a). 爱沙尼亚语句法标注与时间语义标注的整合研究 // 第九届国际语言资源与评估会议(LREC'14)论文集。 2. S.Orasmaa(2014b). 显性时间线索的可获得性对人工时间关系标注的影响 // 《人类语言技术——波罗的海视角》,第215-218页,IOS出版社。



