遇见数据集

Gujarat-HWD: A Dataset of Gujarati Handwritten Digits

收藏
NIAID Data Ecosystem2026-05-02 收录
官方服务:

资源简介:

The Gujarat-HWD dataset offers a comprehensive collection of 11,000 images of handwritten digits in the Gujarati language, representing numerals from 0 to 9. This dataset is primarily designed for applications in optical character recognition (OCR), machine learning, and deep learning, with a particular focus on regional language processing. Gujarati, one of the most widely spoken languages in India, uses a script distinct from Devanagari. However, despite its extensive use, there is a notable lack of publicly available datasets for Gujarati handwritten digits. The Gujarat-HWD dataset bridges this gap by providing a clean, labelled, and diverse set of images that can aid researchers and developers in building effective recognition models for regional scripts. The dataset has been developed through a systematic process involving the collection, scanning, and preprocessing of handwritten digit samples provided by more than 350 individuals from various age groups and educational backgrounds. It is well-suited for training and evaluating classification models using standard convolutional neural network (CNN) architectures. Additionally, the dataset can be extended for use in cross-lingual digit recognition, handwriting analysis, and other regional OCR systems.

创建时间:
2025-07-21
二维码
社区交流群
二维码
科研交流群
商业服务