The CLASSLA-Stanza model for morphosyntactic annotation of non-standard Croatian 2.1
收藏官方服务:
资源简介:
The model for morphosyntactic annotation of non-standard Croatian was built with the [CLASSLA-Stanza tool](https://github.com/clarinsi/classla) by training on the [hr500k training corpus](http://hdl.handle.net/11356/1792) and the [ReLDI-NormTagNER-hr corpus](http://hdl.handle.net/11356/1793), using the [CLARIN.SI-embed.hr word embeddings](http://hdl.handle.net/11356/1790). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~92.49. The model is available for download from the CLARIN.SI repository.
创建时间:
2025-07-04



