遇见数据集

Data for the Journal Paper Machine Learning-Based Context-Aware Lemmatization for Low-Resource Languages: A Case Study of Setswana

收藏
Figshare2025-10-20 更新2026-04-08 收录
官方服务:

资源简介:

Efficient natural language processing (NLP) tools for Setswana are essential for improving human-machine interaction, yet the language remains underrepresented in computational linguistics due to its complex morphology and limited linguistic resources. This study introduces a context-aware machine-learning-based lemmatization model for Setswana, addressing challenges in word sense disambiguation and morphological analysis. Unlike previous rule-based lemmatizers, which process words in isolation, this model incorporates contextual information using Naïve Bayes (NB) and N-gram embeddings to improve lemma prediction accuracy. The proposed model was trained and evaluated using a manually annotated Setswana corpus, integrating part-of-speech (POS) tagging and named entity recognition (NER) as key linguistic features. Performance evaluation, based on accuracy (70.32%), precision (70%), recall (65%), and F1-score (66%), demonstrates the model’s effectiveness in resolving polysemous words, a challenge not addressed by existing Setswana lemmatization approaches. Comparative analysis with prior studies highlights that machine-learning models outperform rule-based approaches in capturing contextual dependencies, although dataset size and feature selection remain critical to performance improvement. This research marks a significant advancement in Setswana NLP, establishing a foundation for future hybrid models that integrate deep learning and rule-based techniques for enhanced accuracy. The study contributes to the development of computational tools for low-resource languages, paving the way for their inclusion in modern information retrieval, machine translation, and conversational AI systems.

提供机构:
Zlotnikova, Irina
创建时间:
2025-10-20
二维码
社区交流群
二维码
科研交流群
商业服务