遇见数据集

A Benchmark Dataset for Manipuri Meetei-Mayek Handwritten Character Recognition

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

A benchmark dataset is always required for any classification or recognition system. To the best of our knowledge, no benchmark dataset exists for handwritten character recognition of Manipuri Meetei-Mayek script in public domain so far. Manipuri, also referred to as Meeteilon or sometimes Meiteilon, is a Sino-Tibetan language and also one of the Eight Scheduled languages of Indian Constitution. It is the official language and lingua franca of the southeastern Himalayan state of Manipur, in northeastern India. This language is also used by a significant number of people as their communicating language over the north-east India, and some parts of Bangladesh and Myanmar. It is the most widely spoken language in Northeast India after Bengali and Assamese languages. In this work, we introduce a handwritten Manipuri Meetei-Mayek character dataset which consists of more than 5000 data samples which were collected from a diverse population group that belongs to different age groups (from 4 years to 60 years), genders, educational backgrounds, occupations, communities from three different districts of Manipur, India (Imphal East District, Thoubal District and Kangpokpi District) during March and April 2019. Each individual was asked to write down all the Manipuri characters on one A4-size paper. The recorded responses are scanned with the help of a scanner and then each character is manually segmented from the scanned images. This dataset consists of segmented scanned images of handwritten Manipuri Meetei-Mayek characters (Mapi Mayek, Lonsum Mayek, Cheitap Mayek, Cheising Mayek, Khutam Mayek) of size 128X128 pixels in .JPG format as well as in. MAT format.

任何分类或识别系统均需配套基准数据集。据我们所知,目前公共领域中尚无针对曼尼普尔语梅泰-梅耶克(Manipuri Meetei-Mayek)手写字符识别的基准数据集。 曼尼普尔语,又称梅泰隆(Meeteilon)或梅泰隆(Meiteilon),属于汉藏语系,同时也是印度宪法第八附表所列的八种官方语言之一。该语言是印度东北部喜马拉雅山南麓曼尼普尔邦的官方语言及通用语,在印度东北部地区、孟加拉国部分区域及缅甸均有大量使用者,是印度东北部地区使用范围最广的语言,仅次于孟加拉语与阿萨姆语。 本研究构建了一套曼尼普尔语梅泰-梅耶克手写字符数据集,共包含5000余条数据样本。样本采集于2019年3月至4月,从印度曼尼普尔邦的三个区县(因帕尔东区、图巴尔区、康波克皮区)的多样化人群中收集,覆盖不同年龄层(4岁至60岁)、性别、教育背景、职业及社群。每位受试者被要求在一张A4纸上书写全部曼尼普尔语字符,采集到手写原稿后通过扫描仪进行扫描,随后由人工对扫描图像中的每个字符进行分割。 本数据集包含分割后的曼尼普尔语梅泰-梅耶克手写字符扫描图像,涵盖马皮梅耶克(Mapi Mayek)、隆桑梅耶克(Lonsum Mayek)、切塔普梅耶克(Cheitap Mayek)、切辛梅耶克(Cheising Mayek)、库塔姆梅耶克(Khutam Mayek)五类字符,图像分辨率为128×128像素,格式包含JPG及MAT两种。

创建时间:
2019-09-27
二维码
社区交流群
二维码
科研交流群
商业服务