kavinh07/ocr_dataset_shamadhan_synth_200k
收藏资源简介:
NID OCR数据集是一个用于图像到文本任务的OCR数据集,专门针对孟加拉语和英语的国民身份证(NID)字段识别。数据集包含合成和真实两种来源的图像:合成数据通过TextRecognitionDataGenerator生成,平衡了文本标记;真实数据来自Shamadhan的NID字段裁剪图像,并经过标签审核。数据集中包括145,800个训练样本和35,363个验证样本。每行数据包含图像(RGB格式的裁剪NID字段)、OCR地面真实文本标签、NID字段类型(如en_name、address_line_00)以及数据来源(synthetic或shamadhan)。该数据集旨在支持OCR模型训练和评估,特别是在多语言NID文档处理场景中。
The NID OCR Dataset is an OCR dataset for image-to-text tasks, specifically designed for recognizing National ID (NID) fields in Bengali and English. It includes images from both synthetic and real sources: synthetic data is generated using TextRecognitionDataGenerator with token balancing, while real data consists of cropped NID field images from Shamadhan with reviewed labels. The dataset contains 145,800 training samples and 35,363 validation samples. Each row includes an image (cropped NID field in RGB format), the OCR ground-truth text label, the NID field type (e.g., en_name, address_line_00), and the data source (synthetic or shamadhan). This dataset is intended for training and evaluating OCR models, particularly in multilingual NID document processing scenarios.



