A Large Labeled Dataset for Complex Layout Analysis of Medieval Manuscripts
收藏官方服务:
资源简介:
This dataset is designed for historical document layout analysis and semantic segmentation. It contains 1,800 pages of medieval manuscript images with corresponding ground truth annotations. The annotations are generated using a semi-automatic iterative process combining automatic segmentation and manual correction. The dataset is split into 1,200 training images, 300 validation images, and 300 test images. Each image is paired with a pixel-level annotation belonging to six non-overlapping classes: background, decoration, filler, text, main body, and drop caps. The dataset is intended for research in document analysis, computer vision, and deep learning.
提供机构:
Zenodo创建时间:
2026-05-19



