A Benchmark Dataset for Printed Meitei/Meetei Script Character Recognition
收藏资源简介:
The Manipuri language is the official language of the Indian state of Manipur. The language belongs to the Tibeto-Burman family of languages. The dataset contains scanned 824 pages of printed documents, along with binarized images, text files, and XML files for each raw image. It also includes 51,460 isolated character samples, composed of 27 consonants, 7 half-consonants, 8 vowels, and 10 numerical. This dataset could be used not only for optical character recognition (OCR) research but also in the different research areas of natural language processing (NLP).
曼尼普尔语(Manipuri)是印度曼尼普尔邦的官方语言,隶属于藏缅语族。本数据集包含824页印刷文档的扫描件,并为每份原始扫描图像配套提供二值化图像、文本文件以及可扩展标记语言(XML)文件。此外,数据集还包含51460个孤立字符样本,由27个辅音、7个半辅音、8个元音与10个数字字符构成。该数据集既可用于光学字符识别(OCR)相关研究,也可应用于自然语言处理(NLP)的多个研究领域。




