遇见数据集

The CLASSLA-Stanza model for lemmatisation of non-standard Serbian 2.1

收藏
SSH Open MarketPlace2023-10-13 更新2024-08-03 收录
官方服务:

资源简介:

The model for lemmatisation of non-standard Serbian was built with the [CLASSLA-Stanza tool](https://github.com/clarinsi/classla) by training on the [SETimes.SR training corpus](http://hdl.handle.net/11356/1200) combined with the [Serbian non-standard training corpus ReLDI-NormTagNER-sr](http://hdl.handle.net/11356/1794) and using the [srLex inflectional lexicon](http://hdl.handle.net/11356/1233). These corpora were additionally augmented for handling missing diacritics by repeating parts of the corpora with diacritics removed. The estimated F1 of the lemma annotations is ~94.92. The model is available for download from the CLARIN.SI repository.

创建时间:
2023-10-13
二维码
社区交流群
二维码
科研交流群
商业服务