遇见数据集

Agapet (Christian Arabic HTR Model) Training Datasets

收藏
Zenodo2025-05-20 更新2026-05-26 收录
官方服务:

资源简介:

These files consist of images and gold-standard, expert-corrected segmentation and transcription of Christian Arabic manuscripts in PAGE .XML format. These images were used for the training and testing of the Agapet HTR models for Christian Arabic hands, available in formats compatible with Transkribus and eScriptorium/Kraken. The specific contents of the dataset are as follows: Full segmentation and transcription of SA-418, a 13th-century manuscript, totalling 342 pages. Full segmentation and transcription of Sinai Arabic 423, a 17th-century manuscript, totalling 478 pages across 239 images 11 segmented and transcribed pages of BnF Arabe 76, a 14th-century manuscript, used for further testing of the eScriptorium models These datasets are released to allow independent training, testing, and development of models for computational recognition and analysis of Christian Arabic hands.

提供机构:
Zenodo
创建时间:
2025-05-20
二维码
社区交流群
二维码
科研交流群
商业服务