Historical Arabic Handwritten Text Recognition Dataset
收藏Mendeley Data2026-04-18 收录
官方服务:
资源简介:
A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image. No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and correction.
本数据集收录了丰富的阿拉伯语历史文本,覆盖数个世纪以来不同地域的语料资源。专家们对四十份历史书页进行了精细转录,其中每五份书页均来自一本独立典籍,为每张扫描图像提供了文本基准真值(ground truth)。 据我们所知,此前尚无同类数据集公开发布。本数据集旨在为关注阿拉伯光学字符识别(Optical Character Recognition,简称OCR)与文本校正的从业者及研究者,提供深度学习OCR建模与测试的支撑资源。
创建时间:
2024-01-23




