Thai Literature Corpora (TLC) 是一个包含泰语古典文学文本的语料库,分为两个数据集:TLC集和TNHC集。TLC集来源于Vajirayana Digital Library,包含按章节和诗节存储的文本,未进行分词处理。TNHC集来源于Thai National Historical Corpus,按行存储并手动分词。数据集支持语言建模和语言生成任务,语言为泰语。数据集的
--- annotations_creators: - no-annotation language_creators: - crowdsourced language: - bo license: - other multilinguality: - monolingual pretty_name: Tibetan Classical Buddhist Text Corpus size_cate
Ontology is a plain text file containing statements in the Turtle syntax forming OpenBiodiv-O. It can be edited in a text (e.g. Sublime Text, Emacs, etc.) or in an ontology editor (e.g. ProtĂŠgĂŠ). It