遇见数据集

Testing Dataset for Rejang Language Spelling Correction Using Hybrid Euclidean Distance and N-Gram Methods

收藏
Zenodo2026-06-20 更新2026-06-21 收录
官方服务:

资源简介:

This dataset supports the experimental evaluation of a hybrid spelling correction framework designed for low-resource languages, specifically focusing on the Coastal dialect of the Rejang language. The dataset contains 1,000 test tokens subjected to keyboard proximity error simulations. It includes comparative performance metrics evaluating the proposed character-level N-gram and Euclidean distance combination against industry-standard benchmarks, namely Levenshtein distance and Jaro-Winkler distance. The evaluation metrics focus on system accuracy, precision, recall, and F1-score to analyze model stability and phonetic variation adaptation under constrained resource scenarios.

提供机构:
Zenodo
创建时间:
2026-06-20
二维码
社区交流群
二维码
科研交流群
商业服务