Statistical analysis of The type of Syllable error in Tibetan text
收藏资源简介:
Tibetan information processing is greatly influenced by the quality of Tibetan texts. It is an effective way to improve the quality of Tibetan texts to correct the complex and diverse syllable errors. This paper uses real Tibetan text as statistical source, which contains more than 150 million syllables. 2333617 wrong syllables were found by computer, accounting for 5.6% of the total corpus texts. According to the context information and Tibetan grammar rules, error syllables were manually corrected and classified into 11 types, and the occurrence frequency and high frequency error syllables of each type were counted. By analyzing the error causes, it provides some references for the design and implementation of Tibetan text proofreading system and other software.



