IncuLay: Incunabula Layout Dataset
收藏资源简介:
We introduce IncuLay (Incunabula Layout Dataset), a dataset of 605 annotated page images derived from incunabula and related early printed books held by the Jagiellonian Digital Library.Each page is manually annotated with bounding boxes corresponding to semantically meaningful layout classes, enabling systematic evaluation of document layout analysis in this domain. IncuLay aims to fill the gap left by DocLayNet and PubLayNet—the two largest datasets for layout analysis—which do not include early printed books. IncuLay's annotations are provided in COCO format and are compatible at a high level with both datasets, though not identical, in order to capture the specific characteristics of incunabula. The repository contains the following files: README.md — description of the dataset data.zip — IncuLay dataset (one COCO file and 605 image files in 6 subfolders) browse_IncuLay.ipynb — simple Jupyter Notebook to load and browse the dataset preparation_scripts.zip — the whole workflow used to create a unified dataset source_data.zip — previously unpublished jdl_incunabula_500 dataset, used as an input for preparation scripts



