Testing Dataset for Rejang Language Spelling Correction Using Hybrid Euclidean Distance and N-Gram Methods
收藏官方服务:
资源简介:
This dataset supports the experimental evaluation of a hybrid spelling correction framework designed for low-resource languages, specifically focusing on the Coastal dialect of the Rejang language. The dataset contains 1,000 test tokens subjected to keyboard proximity error simulations. It includes comparative performance metrics evaluating the proposed character-level N-gram and Euclidean distance combination against industry-standard benchmarks, namely Levenshtein distance and Jaro-Winkler distance. The evaluation metrics focus on system accuracy, precision, recall, and F1-score to analyze model stability and phonetic variation adaptation under constrained resource scenarios.
提供机构:
Zenodo创建时间:
2026-06-20



