IndOCR-24: An Industrial OCR Dataset for Restricted-Alphabet Character Classification
收藏资源简介:
IndOCR-24 is a dataset developed for research on industrial Optical Character Recognition (OCR) involving restricted-alphabet character classification. The dataset was created from real industrial character images extracted from traceability codes printed on copper cathodes under operational conditions. The repository contains the complete set of original industrial character images, representative samples from the augmented training and validation datasets, and the Python scripts required to reproduce the complete augmented dataset used in the associated experiments. The augmentation pipeline applies photometric, geometric, and morphological transformations to the original images in order to increase variability while preserving the original class labels. The dataset is intended for benchmarking OCR algorithms, character classification methods, deep learning models, transfer learning, and data augmentation strategies in industrial computer vision applications.
IndOCR-24是为开展受限字符集分类的工业光学字符识别(Optical Character Recognition, OCR)研究而开发的数据集。该数据集的样本源自实际运行工况下,从铜阴极上打印的追溯码中提取的真实工业字符图像。 本数据集仓库包含完整的原始工业字符图像集、增强训练与验证数据集的代表性样本,以及复现相关实验所用完整增强数据集所需的Python脚本。该增强流水线会对原始图像施加光度、几何与形态学变换,在保留原始类别标签的前提下提升数据多样性。 本数据集旨在为工业计算机视觉应用中的OCR算法、字符分类方法、深度学习模型、迁移学习以及数据增强策略提供基准测试支持。




