官方服务:
资源简介:
A Comprehensive Resource for Sindhi Language Processing
关于信德语处理的一个全面资源
应用场景:
创建时间:
2024-07-27
相关数据集
yalhessi/lemexp
该数据集包含两个配置:default和hol。每个配置都包含三个主要特征:defs(定义)、symbols(符号)和lemma(词元)。数据集主要用于训练,包含32920个示例。default配置的总数据大小为75847048字节,而hol配置的总数据大小为73239698字节。这些特征表明数据集可能用于自然语言处理任务,如词义消歧或语义分析。
Hugging Face2024-07-02 更新130
The CLASSLA-Stanza model for lemmatisation of non-standard Serbian 2.1
The model for lemmatisation of non-standard Serbian was built with the [CLASSLA-Stanza tool](https://github.com/clarinsi/classla) by training on the [SETimes.SR training corpus](http://hdl.handle.net/
SSH Open MarketPlace2023-10-13 更新40
Lemma list of the Beseda Corpus Lemmatisation Lexicon (ELEXIS)
Lematizacijski slovar (leksikon besednih oblik za Besedo). Beseda Corpus Lemmatisation Lexicon for Slovenian language was generated at the Fran Ramovš Institute of Slovenian...
B2FIND20
nikopartanen/old-literary-finnish-lemmatization: Old Literary Finnish Lemmatization Dataset
This is a dataset that contains randomly selected and manually lemmatized sentences from the corpus of Old Literary Finnish. Please cite and consult the original corpus as well: Institute for the
NIAID Data Ecosystem60
Data for the Journal Paper Machine Learning-Based Context-Aware Lemmatization for Low-Resource Languages: A Case Study of Setswana
Efficient natural language processing (NLP) tools for Setswana are essential for improving human-machine interaction, yet the language remains underrepresented in computational linguistics due to its
DataCite Commons2025-10-20 更新30



