遇见数据集

uTHCD

收藏
arXiv2021-03-13 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

uTHCD是一个全面的非约束性泰米尔手写文字数据库,包含约91000个样本,覆盖156个字符类别,每个类别约600个样本。数据集由阿曼技术大学信息技术系创建,通过邀请志愿者在特定网格内书写样本来收集离线样本,而在线样本则通过数字书写板收集。数据集涵盖了广泛的书写风格和由于离线扫描过程产生的固有畸变,如笔划不连续性和笔划厚度变化等。uTHCD旨在为泰米尔手写文字识别研究提供新的基准,并作为文档图像分析领域多个研究方向的起点。

uTHCD is a comprehensive unconstrained Tamil handwritten character database comprising approximately 91,000 samples spanning 156 character categories, with around 600 samples per category. This dataset was created by the Department of Information Technology, Oman University of Technology. Offline samples were collected by inviting volunteers to write characters within a specified grid, while online samples were gathered using digital writing tablets. The dataset covers a wide range of writing styles and inherent distortions caused by the offline scanning process, such as stroke discontinuities and variations in stroke thickness. uTHCD aims to provide a new benchmark for Tamil handwritten character recognition research and serve as a starting point for multiple research directions in the field of document image analysis.

创建时间:
2021-03-13
二维码
社区交流群
二维码
科研交流群
商业服务