遇见数据集

Annotated Corpora of Historical Catalan (HisCat) - Llibre dels Fets

收藏
Zenodo2021-10-29 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

This repository is part of the Annotated Corpora of Historical Catalan (HisCat). It contains the first POS-tagged text that is partially manually corrected and used to train Old Catalan POS taggers, described in the following paper: Meelen, Marieke &amp; Pujol i Campeny, Afra, (2021) 'Old Catalan Morphosyntax: developing an annotated corpus' in <em>Journal of Open Humanities Data</em>. This POS-tagged text is the 13th century <em>Llibre dels Fets</em>, a historical chronicle. The version of the text used for this project is Bruguera, J. (1991). <em>El Llibre dels Fets del Rei en Jaume</em>. Barcelona: Barcino. as prepared for the <em>Corpus Informatitzat del Català Antic</em> Torruella, J., Pérez Saldanya, M., &amp; Martines, J. (2009). <em>Corpus Informatitzat del Català Antic</em>. URL: http://cica.cat/. The subcorpus counts with 164,096 POS-annotated tokens (165,538 tokens including punctuation and folio markers), of which 60,000 have been manually corrected. This subcorpus contains a total of and 4,506 main clauses. POS tagging of this text was done with the Memory-Based Tagger by TiMBL (https://languagemachines.github.io/mbt/). The code accompanying the paper can be found on GitHub: https://github.com/lothelanor/catalancorpora). In addition to memory-based tagging, have tried neural-based tagging with TARGER (https://github.com/achernodub/targer) for which we created word embeddings that can be found on Zenodo. Results for memory-based tagging were better, however, which is why this version is uploaded here. <pre> </pre>

创建时间:
2021-10-29
二维码
社区交流群
二维码
科研交流群
商业服务