遇见数据集

A Vast Dataset for Kurdish Digits and Isolated Characters Recognition

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Kurdish language dialects are used across four main nation-states in the Middle East, and only one dialect, Sorani, has official status in one of these nation-states. The majority of Kurdish-speaking regions are located in Turkey, Iraq, Iran, and Syria. More than 30 million people speak Kurdish as a whole, according to estimates. One of the two main dialects of Kurdish, known as Central Kurdish (Sorani), is spoken by an estimated 9 to 10 million people. It is mostly written with a 35-character modified Arabic/Persian alphabet and includes characters that have recently been replaced, such as (ك) which is no longer used by the Kurdish language and has been replaced with (ک). This work presents two massive datasets for central Kurdish handwriting digits and isolated characters named K-ZHMARA and K-PIT. The first dataset, named K-ZHMARA dataset, contains 70,000 images of Kurdish digits, 7,000 images for each digit, and a printed A4 paper with a grid of 10 × 10 is used for data collection. Apart from digits, the K-PIT dataset includes 245,000 images of all Kurdish characters, 7,000 images for each character; data was collected via a printed A4 paper with a grid of 12 × 10 for this dataset. Moreover, both datasets include 315,000 images. Then, using Python programming, each piece of paper was scanned, segmented, cropped, resized, binarized, and inverted using edge detection and image processing techniques. Most students from the University of Halabja and the primary and preparatory school in the Halabja governorate volunteered to fill out the forms. Furthermore, these datasets are suitable for Kurdish isolate handwritten optical digit/character recognition. Labeling and organizing: Each image is labeled with an ID number, the number of the folder in each dataset represents a single digit or character. For example, folder number 02 in the K-PIT dataset is the id of the letter, which in this case is Alef (ا), and folder number 03 in the K-ZHMARA dataset is the id of the digit, which in this case is three (٣). Each digit and character were stored in a folder with its ID as the name of that folder, with each folder containing 6000 images of that letter/digit for the training and 1000 images for the testing.

库尔德语方言分布于中东四个主要主权国家境内,仅索拉尼(Sorani)一种方言在其中一个国家拥有官方地位。库尔德语使用者主要聚居在土耳其、伊拉克、伊朗及叙利亚境内。据估算,全球使用库尔德语的总人口超过3000万。 库尔德语两大主要方言之一的中央库尔德语(Central Kurdish,索拉尼),使用者约为900万至1000万。该方言主要采用35个字符的改良阿拉伯/波斯字母表书写,其中部分字符已被弃用,例如原字符(ك)现已不再被库尔德语使用,替换为(ک)。 本研究构建了两套针对中央库尔德语手写数字与孤立字符的大规模数据集,分别命名为K-ZHMARA与K-PIT。首个数据集K-ZHMARA包含7万张库尔德语手写数字图像,每个数字对应7000张样本,数据采集依托带有10×10网格的打印A4纸张完成。除数字外,K-PIT数据集涵盖全部库尔德语字符的24.5万张手写图像,每个字符对应7000张样本,该数据集的数据采集依托带有12×10网格的打印A4纸张完成。两套数据集总计包含31.5万张图像。随后,研究团队通过Python编程,对每张采集纸张进行扫描、分割、裁剪、缩放、二值化处理,并结合边缘检测与图像处理技术完成图像反转。哈尔卜贾大学(University of Halabja)以及哈尔卜贾省小学与预科学校的多数学生志愿参与了手写表单填写工作。此外,这两套数据集适用于库尔德语孤立手写数字/字符的光学识别任务。 标注与组织规则: 每张图像均以ID编号进行标注,每个数据集内的文件夹编号对应单个数字或字符的ID。例如,K-PIT数据集中编号为02的文件夹对应字母阿列夫(ا)的ID,而K-ZHMARA数据集中编号为03的文件夹对应数字3(٣)的ID。每个数字与字符均以其ID作为文件夹名称进行存储,每个文件夹包含该字母/数字的6000张训练样本与1000张测试样本。

创建时间:
2022-12-22
二维码
社区交流群
二维码
科研交流群
商业服务