遇见数据集

Statistical analysis of The type of Syllable error in Tibetan text

收藏
科学数据银行2021-12-10 更新2026-04-23 收录
官方服务:

资源简介:

Tibetan information processing is greatly influenced by the quality of Tibetan texts. It is an effective way to improve the quality of Tibetan texts to correct the complex and diverse syllable errors. This paper uses real Tibetan text as statistical source, which contains more than 150 million syllables. 2333617 wrong syllables were found by computer, accounting for 5.6% of the total corpus texts. According to the context information and Tibetan grammar rules, error syllables were manually corrected and classified into 11 types, and the occurrence frequency and high frequency error syllables of each type were counted. By analyzing the error causes, it provides some references for the design and implementation of Tibetan text proofreading system and other software.

创建时间:
2021-12-08
二维码
社区交流群
二维码
科研交流群
商业服务