遇见数据集

The CLASSLA-Stanza model for morphosyntactic annotation of non-standard Slovenian 2.1

收藏
SSH Open MarketPlace2025-07-04 更新2025-07-05 收录
官方服务:

资源简介:

The model for morphosyntactic annotation of non-standard Slovenian was built with the [CLASSLA-Stanza tool](https://github.com/clarinsi/classla) by training on the [SUK training corpus](http://hdl.handle.net/11356/1747) and on the [Janes-Tag corpus](http://hdl.handle.net/11356/1732) using the [CLARIN.SI-embed.sl word embeddings](http://hdl.handle.net/11356/1204) expanded with the [MaCoCu-sl Slovene web corpus](http://hdl.handle.net/11356/1517). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~92.17. The model is available for download from the CLARIN.SI repository.

创建时间:
2025-07-04
二维码
社区交流群
二维码
科研交流群
商业服务