musnad-letter-classification
收藏资源简介:
Musnad字母分类数据集是一个用于古南阿拉伯语(也门Musnad)字母识别的合成图像数据集,专为图像分类任务设计。它包含32个不同的Musnad字母类别,每个类别提供1,000张图像,总计32,000张图像。所有图像均为96×96像素的灰度PNG格式。数据集通过使用Noto Sans Old South Arabian字体渲染Unicode Musnad字符生成,并应用了多种数据增强技术以提高模型鲁棒性,包括旋转、平移、高斯模糊、噪声注入、字体大小变化和笔画粗细变化。预期应用场景涵盖Musnad光学字符识别(OCR)研究、古文字识别、图像分类模型开发、文化遗产人工智能项目以及低资源语言技术探索。需要注意的是,该数据集为合成数据,不包含真实石刻铭文中的雕刻变化、风化效果、光照条件或考古文物特征,因此在实际应用时需考虑其与真实场景的差异。基准实验表明,使用该数据集训练的卷积神经网络在验证集上取得了超过99.8%的准确率。
The Musnad Alphabet Classification Dataset is a synthetic image dataset for ancient South Arabian (Yemeni Musnad) alphabet recognition. It is designed for image classification tasks and includes 32 different Musnad letter categories, with 1,000 images per category, totaling 32,000 images. All images are in 96×96 pixel grayscale PNG format. The dataset is generated by rendering Unicode Musnad characters using the Noto Sans Old South Arabian font and applies various data augmentation techniques to enhance model robustness, including rotation, translation, Gaussian blur, noise injection, font size variation, and stroke thickness variation. Intended applications include Musnad optical character recognition (OCR) research, ancient script recognition, image classification model development, cultural heritage AI projects, and low-resource language technology exploration. It is important to note that this dataset is synthetic and does not include real-world variations such as carving effects, weathering, lighting conditions, or archaeological artifact features, so differences from real-world scenarios should be considered in practical use. Benchmark experiments show that convolutional neural networks trained on this dataset achieve over 99.8% accuracy on the validation set.
数据集名称
Musnad 字母分类数据集(Musnad Letter Classification Dataset)
数据集概述
- 任务类型:图像分类
- 语言:阿拉伯语
- 类别数:32 个 Musnad 字母类别
- 图像总数:32,000 张
- 每类图像数:1,000 张
- 图像尺寸:96×96 像素
- 图像格式:灰度 PNG 图像
数据生成方法
- 使用 Unicode Musnad 字符,以 Noto Sans Old South Arabian 字体渲染生成。
- 数据增强技术包括:
- 旋转
- 平移
- 高斯模糊
- 噪声注入
- 字体大小变化
- 笔画粗细变化
基线 CNN 结果
- 类别数:32
- 总图像数:32,000
- 训练图像数:25,600
- 验证图像数:6,400
- 图像尺寸:96×96 灰度
- 最佳验证准确率:99.87%
- 最终验证准确率:99.84%
预期用途
- Musnad OCR 研究
- 古文字识别
- 图像分类
- 文化遗产人工智能
- 低资源语言技术
局限性
该数据集为合成数据集,不代表真实的石刻铭文、雕刻变化、风化、光照条件或考古文物。




