遇见数据集
官方服务:

资源简介:

This is a high quality research benchmark dataset for handwritten Bengali (Bangla) digit recognition. It consists of two structurally consistent variants of 20,000 and 60,000 samples respectively (BanglaDigit-20k and BanglaDigit-60k). To prevent compression artifacts, the data was obtained from 10 native Bengali speakers through an independent session on a tablet interface interface with a capacitive stylus, then saved as an uncompressed PDF. Unlike the data sets that use a geometric bounding-box centering method, BanglaDigit adopts Intensity Center of Mass (CoM) alignment technique, which is a mathematically sound curation methodology aligning each numeral in accordance with the weighted pixel intensity distribution to remove the translation bias systematically embedded in the Bengali script. All of the images are gray scale png and normalized to 28×28 pixels. All the digits (০–৯) of Bengali language have been taken with 2000 base samples for each class in the 20k version of the dataset, and 6000 samples per each class in the 60k version, using a stochastic augmentation engine. This is a dataset designed to provide a repeatable and stringent standard for both optical character recognition (OCR) research in the South Asian region and for quick architectural prototyping and testing, and deployment of edge computing. The six CNN architectures tested, ShuffleNetV2, MobileNetV2, ResNet18, DenseNet121, EfficientNetB0 and WideResNet50 achieve state of the art accuracy of 99.95% and 99.86% on the 20k and 60k splits respectively with a fixed training budget of 10 epochs.

本数据集为手写孟加拉文(孟加拉语)数字识别领域的高质量研究基准数据集。它包含两个结构一致的子集,样本规模分别为20000与60000,分别命名为BanglaDigit-20k与BanglaDigit-60k。为避免压缩伪影,该数据集的数据由10名母语为孟加拉语的受访者通过平板界面配合电容触控笔独立采集,随后保存为未压缩PDF格式。与采用几何边界框居中方法的数据集不同,BanglaDigit采用了强度质心(Intensity Center of Mass, CoM)对齐技术,这是一种严谨的数学整理方法,可根据加权像素强度分布对每个数字进行对齐,系统性消除孟加拉文脚本中固有的平移偏差。所有图像均为灰度png格式,并统一归一化至28×28像素尺寸。该数据集的20k版本中,孟加拉文所有数字(০–৯)的每个类别均包含2000个基础样本;60k版本中每个类别则包含6000个样本,样本生成过程使用了随机数据增强引擎。本数据集旨在为南亚地区的光学字符识别(Optical Character Recognition, OCR)研究提供可复现且严苛的标准基准,同时可用于快速的模型架构原型开发、测试以及边缘计算部署。在固定10个训练轮次(epoch)的训练预算下,本次测试的六种卷积神经网络(Convolutional Neural Network, CNN)架构——ShuffleNetV2、MobileNetV2、ResNet18、DenseNet121、EfficientNetB0及WideResNet50——在20k与60k划分的测试集上分别取得了99.95%与99.86%的当前最优准确率。

创建时间:
2026-06-11
二维码
社区交流群
二维码
科研交流群
商业服务