遇见数据集

Tamil Palm Leaf Character Dataset: Single and Multi-Segmented Characters Across 20 Different Centuries.

收藏
DataCite Commons2025-04-01 更新2025-04-16 收录
官方服务:

资源简介:

This dataset consists of Tamil palm leaf manuscript Characters collected from 20 different centuries. The dataset is divided into single-segmented characters (individual Tamil letters) and multi-segmented characters (complex characters composed of multiple strokes). The single segmented characters are labelled from label 1 to label 97. Label 6 and label 20 are the most appeared characters where both characters are used in different centuries with slight variations. Label 89 is less appeared and it is not much used, whereas in multi-segmented the characters are labelled from label 98 to label 1045.Most of them are appeared only once and they are fully composed of multi strokes. Each label in the dataset represents different pixel properties extracted from the corresponding images. With 1045 labels, the dataset captures a diverse range of character variations. The images are of size 40 × 32 pixels, and the pixel properties are averaged to provide meaningful feature representations. It is designed to aid research in historical manuscript recognition, optical character recognition (OCR), and Tamil script analysis. The dataset includes high-resolution binarized images along with annotations that provide segmentation information. The images are pre-processed to enhance clarity, making them suitable for deep learning applications in character recognition. It is an unbalanced dataset and it requires augmentation and sampling for the better accuracy. The folder structure inside are the ZIP file format. Files & Data Organization: • Images: High-resolution palm leaf manuscript images (.PNG format). • Annotations: Image files containing ground truth text are manually annotated and segmentation details. • Single Characters: Segmented images of individual Tamil Characters. • Multi Segmented Characters: Segmented images of Composite Tamil characters. Methodology: • Data sourced from Tamil palm leaf manuscripts from Tamil Virtual Academy (TVA). • Preprocessing steps include noise reduction and Binarization. • Characters are Semantically segmented and categorized based on the ground truth. Ethical Considerations: This dataset contains historical materials and it should be used with proper attribution. The dataset is shared for research purposes only, and modifications that misrepresent its content are not permitted. Funding & Acknowledgments: This dataset is developed by SASTRA Deemed University with the support of Tamil Virtual Academy (TVA) for providing access to manuscript images.

本数据集包含来自20个不同历史时期的泰米尔棕榈叶手稿字符。本数据集分为单分割字符(单个泰米尔字母)与多分割字符(由多笔画构成的复合字符)两类。单分割字符的标签范围为1至97。标签6与标签20为出现频次最高的字符,二者在不同历史时期均有使用且存在细微变体。标签89的出现频次最低且使用场景极少;多分割字符的标签范围则为98至1045,其中绝大多数仅出现一次,且均由多笔画构成。数据集中的每个标签对应从对应图像中提取的不同像素属性。本数据集共包含1045个标签,涵盖了丰富多样的字符变体。图像分辨率为40×32像素,通过对像素属性取平均以生成具有实际意义的特征表示。本数据集旨在助力历史手稿识别、光学字符识别(Optical Character Recognition, OCR)以及泰米尔文字分析领域的研究工作。数据集包含高分辨率二值化图像以及提供分割信息的标注文件。图像已经过预处理以提升清晰度,适用于字符识别相关的深度学习应用。本数据集存在类别不平衡问题,需通过数据增强与采样手段提升模型识别精度。数据集内部的文件夹结构采用ZIP压缩格式。 文件与数据组织: • 图像:高分辨率棕榈叶手稿图像(格式为.PNG)。 • 标注:包含真值文本的图像文件均经过手动标注,并附带分割相关细节。 • 单字符:单个泰米尔字符的分割图像。 • 多分割字符:复合泰米尔字符的分割图像。 研究方法: • 数据来源:数据取自泰米尔虚拟学院(Tamil Virtual Academy, TVA)提供的泰米尔棕榈叶手稿。 • 预处理流程:包含降噪与二值化处理步骤。 • 字符处理:基于真值文本对字符进行语义分割与分类。 伦理考量: 本数据集包含历史文物资料,使用时需注明正确来源。本数据集仅用于学术研究用途,严禁对其内容进行篡改或歪曲。 资助与致谢: 本数据集由SASTRA Deemed University开发,并得到泰米尔虚拟学院(Tamil Virtual Academy, TVA)的支持,后者提供了手稿图像的访问权限。

提供机构:
Mendeley Data
创建时间:
2025-03-24
搜集汇总
数据集介绍
Tamil Palm Leaf Character Dataset: Single and Multi-Segmented Characters Across 20 Different Centuries. 数据集图片
背景与挑战
背景概述
该数据集是一个泰米尔语棕榈叶手稿字符集合,覆盖20个不同世纪,包含单分割字符(97个标签)和多分割字符(总计1045个标签),图像尺寸统一为40×32像素并经过预处理。它专为历史手稿识别、光学字符识别(OCR)和泰米尔文字分析研究设计,但属于不平衡数据集,需数据增强和采样以提高模型准确性。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务