kailabresearch/khld-kurdish-handwritten-sentences
收藏资源简介:
库尔德手写行数据集(KHLD)是一个大规模基准数据集,专为中央库尔德语(索拉尼语)的手写文本识别(HTR)和光学字符识别(OCR)系统设计。该数据集包含47,944张手写行图像,组织成4,802个独特句子,每个句子有10个变体。数据集捕捉了库尔德字母(基于阿拉伯文字并带有修改)的独特正字法特征,包括特定字符如ڤ、ڕ、ڵ、ێ和ۆ。总图像数为47,944张,独特句子数为4,802个,每个句子的变体数为10个,格式为JPEG(灰度,已调整大小),原始扫描分辨率为600-1,200 DPI。数据集结构按文件夹组织,每个文件夹对应一个独特的库尔德句子,包含10个图像文件,代表该句子的不同手写样本。数据字段包括:image(手写行图像)、text(真实键入的库尔德句子,索拉尼语)、folder_id(句子组的ID)和variation_id(特定手写变体的ID)。数据集的创建是为了填补AI和OCR领域中库尔德语(索拉尼语)作为低资源语言的空白,确保高质量、大规模基准的可用性。数据来源包括库尔德网站、书籍、文章和社交媒体,以混合正式和口语语言风格。数据收集和处理涉及来自伊拉克库尔德斯坦地区埃尔比勒大学和研究所的母语索拉尼语使用者(年龄15-55岁),使用微调的YOLOv8模型(准确率95%以上)和OpenCV模板匹配进行行分割,图像转换为JPEG、灰度化并调整大小以确保一致性。
The Kurdish Handwritten Lines Dataset (KHLD) is a large-scale benchmark designed for handwritten text recognition (HTR) and optical character recognition (OCR) systems specifically for Central Kurdish (Sorani). It consists of 47,944 handwritten line images organized into 4,802 unique sentences, with 10 variations each. The dataset captures the unique orthographic features of the Kurdish alphabet (based on Arabic script with modifications), including specific characters such as ڤ, ڕ, ڵ, ێ, and ۆ. Total Images: 47,944, Unique Sentences: 4,802, Variations per Sentence: 10, Format: JPEG (Grayscale, resized), Resolution: 600–1,200 DPI (original scan). The dataset is organized into folders where each folder corresponds to a unique Kurdish sentence, each folder contains 10 image files representing different handwriting samples of that sentence. Data fields include: image (the handwritten line image), text (the ground truth typed Kurdish sentence, Sorani), folder_id (the ID of the sentence group), and variation_id (the ID of the specific handwriting variation). The dataset was created to fill the gap in high-quality, large-scale benchmarks for Kurdish handwritten text recognition, as Kurdish (Sorani) is a low-resource language in AI and OCR. Source data includes Kurdish websites, books, articles, and social media to ensure a mix of formal and colloquial language styles. Data collection and processing involved native Sorani speakers (ages 15–55) at universities and institutes in Erbil, Kurdistan Region, Iraq, with line segmentation using a fine-tuned YOLOv8 model (achieving 95%+ accuracy) and OpenCV template matching, and preprocessing with images converted to JPEG, grayscaled, and resized for consistency.



