A Benchmark Dataset for Printed Meitei/Meetei Script Character Recognition

NIAID Data Ecosystem2026-03-13 收录

下载链接：

https://data.mendeley.com/datasets/rw4b2zdk95

下载链接

链接失效反馈

官方服务：

资源简介：

The Manipuri language is the official language of the Indian state of Manipur. The language belongs to the Tibeto-Burman family of languages. The dataset contains scanned 824 pages of printed documents, along with binarized images, text files, and XML files for each raw image. It also includes 51,460 isolated character samples, composed of 27 consonants, 7 half-consonants, 8 vowels, and 10 numerical. This dataset could be used not only for optical character recognition (OCR) research but also in the different research areas of natural language processing (NLP).

创建时间：

2022-04-29

5,000+

优质数据集

54 个

任务类型

进入经典数据集