five

Gurmukhi dataset

收藏
NIAID Data Ecosystem2026-05-02 收录
下载链接:
https://data.mendeley.com/datasets/h65gdk4ptv
下载链接
链接失效反馈
官方服务:
资源简介:
This dataset comprises a meticulously augmented collection of Gurmukhi handwritten characters, designed to enhance the performance of machine learning models in optical character recognition (OCR) and related tasks. It includes characters across 41 distinct classes, each augmented to reach a total of approximately 290 samples per class. Key Features: Gurmukhi Script Focus: The dataset exclusively features handwritten characters from the Gurmukhi script, catering specifically to applications involving Punjabi language processing. Diverse Augmentations: Images have been subjected to a range of transformations, including rotations, shifts, shears, zooms, and horizontal flips, promoting robustness to variations encountered in handwritten text. Consistent Dimensions: All images are resized to a uniform 256x256 resolution, ensuring compatibility with most deep learning architectures. Class-Specific Organization: Images are neatly organized into 41 folders, each representing a distinct Gurmukhi character, facilitating targeted training and evaluation. Handwritten Data Collection: The original images used for augmentation were collected from 10 volunteers, introducing natural variability in writing styles and further enhancing the dataset's diversity. Potential Use Cases: Gurmukhi OCR: Train and evaluate OCR models specifically for Gurmukhi script recognition. Handwriting Recognition: Develop models capable of recognizing and transcribing handwritten Gurmukhi text. Script Style Analysis: Explore the variations in handwriting styles within the Gurmukhi script.
创建时间:
2024-09-24
5,000+
优质数据集
54 个
任务类型
进入经典数据集
二维码
社区交流群

面向社区/商业的数据集话题

二维码
科研交流群

面向高校/科研机构的开源数据集话题

数据驱动未来

携手共赢发展

商业合作