The Annotated Corpus of Classical Tibetan (ACTib), Part II - POS-tagged version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL
收藏资源简介:
This corpus is a part-of-speech tagged version of Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, & Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218 using the training data of Hill, Nathan W., & Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878 Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already. using the memory based tagger of https://languagemachines.github.io/mbt/
本语料库(corpus)是对Wallman、Jeff、Rowinski、Zach、Ngawang Trinley、Tomlinson、Chris与Keutzer, Kurt(2017年)发布的《由佛教数字资源中心(Buddhist Digital Resource Center)汇编的藏语电子文本集》[数据集(Data set)](Zenodo,http://doi.org/10.5281/zenodo.821218)进行词性标注(part-of-speech tagged)后得到的版本,其标注所用训练数据源自Hill、Nathan W.与Garrett、Edward(2017年)发布的《古典藏语词性标注(part-of-speech,简称POS)语料库》[数据集(Data set)](Zenodo,http://doi.org/10.5281/zenodo.574878)。 请注意:本数据集的所有文件均未经过后处理或人工校正,且KarmaDelek目录下的少量文件虽原始XML输入已损坏,但仍完成了标注。 本语料库采用了https://languagemachines.github.io/mbt/ 提供的基于记忆的标注器(memory based tagger)。



