THE TOKENIZATION PREMIUM: MEASURING UZBEK TOKENIZATION OVERHEAD IN LARGE LANGUAGE MODELS AND EVALUATING THE 2026 ALPHABET REFORM
收藏资源简介:
Large language models process text as tokens, making tokenization efficiency important for computational cost and context capacity. This study examines Uzbek tokenization efficiency relative to English and other languages and evaluates whether the proposed 2026 Uzbek alphabet reform can reduce the observed overhead. Six widely used tokenizers are compared using a multilingual parallel corpus and a large Uzbek literary corpus. Latin Uzbek requires substantially more tokens than English across the evaluated tokenizers, while Cyrillic Uzbek shows an even greater overhead in some cases. Turkish generally requires fewer tokens than Uzbek under the same tokenizer, suggesting that agglutinative morphology alone does not explain the difference. The proposed alphabet reform produces only modest and tokenizer-dependent reductions in token counts. These findings indicate that orthographic reform can improve tokenization efficiency to a limited extent, but the larger Uzbek tokenization overhead remains primarily a broader tokenizer and language-resource problem.



