Hoda Farsi Digit Dataset
收藏资源简介:
Hoda数据集是第一个手写波斯数字数据集,由Tarbiat Modarres大学的一个硕士项目开发,该项目名为:识别SANJESH注册表格中的波斯数字和字符。该项目与Hoda系统公司合作完成,于2005年夏季在Ehsanollah Kabir教授的监督下完成。数据集样本是从伊朗大学入学考试的约12000份注册表格中提取的手写字符。数据集规格如下:样本分辨率:200 dpi;总样本数:102,352个样本;训练样本:60,000个样本;测试样本:20,000个样本;剩余样本:22,352个样本。每个类别的样本数量不同。
The Hoda dataset is the first handwritten Persian numeral dataset, developed as part of a master's project at Tarbiat Modarres University. The project, titled 'Recognition of Persian Numerals and Characters in SANJESH Registration Forms,' was completed in collaboration with Hoda System Company under the supervision of Professor Ehsanollah Kabir during the summer of 2005. The dataset samples were extracted from approximately 12,000 registration forms of the Iranian university entrance examination. The specifications of the dataset are as follows: sample resolution: 200 dpi; total number of samples: 102,352; training samples: 60,000; test samples: 20,000; remaining samples: 22,352. The number of samples varies across different categories.
Hoda Farsi Digit Dataset 概述
数据集基本信息
- 名称: Hoda Farsi Digit Dataset
- 描述: 该数据集是首个手写波斯数字数据集,由Tarbiat Modarres大学的一个硕士项目开发,项目名称为“Recognizing Farsi Digits and Characters in SANJESH Registration Forms”,与Hoda System Corporation合作完成。
- 开发时间: 2005年夏季
- 监督者: Prof. Ehsanollah Kabir
数据集规格
- 分辨率: 200 dpi
- 总样本数: 102,352
- 训练样本数: 60,000
- 测试样本数: 20,000
- 剩余样本数: 22,352
样本分布
| 数字 | 样本数 |
|---|---|
| 0 | 10070 |
| 1 | 10330 |
| 2 | 9923 |
| 3 | 10334 |
| 4 | 10333 |
| 5 | 10110 |
| 6 | 10254 |
| 7 | 10363 |
| 8 | 10264 |
| 9 | 10371 |
使用许可
- 许可: 免费提供给研究和非商业用途
数据集样本
- 样本多样性: 包含不同书写风格和质量的样本
数据集读取
- 文件格式:
.cdb - 读取代码示例: 提供了Python代码示例,用于读取训练、测试和剩余样本的图像和标签。
数据集网站




