HORAE_Minit. Book of hours, Miniatures and Initials: dataset
收藏资源简介:
HORAE_Minit (Books of Hours, Miniatures and Initials) is a comprehensive annotated dataset for automatic detection and classification of decorative elements in medieval Books of Hours manuscripts. Created at the Institut de Recherche et d'Histoire des Textes (IRHT-CNRS) as part of the Biblissima+ ÉquipEx, this dataset enables large-scale computational analysis of medieval illuminations. Dataset Overview 14,225 images from over 1,000 manuscripts 105,682 annotations across 6 classes of decorative elements 12,270 annotated images with detailed bounding boxes YOLO-compatible format for immediate use with state-of-the-art detection models Complete IIIF traceability with URLs to original digital libraries Detection Classes The dataset includes six classes of decorative elements commonly found in Books of Hours: Miniature (1,186 instances): Narrative scenes located in the text field Historiated initial (496 instances): Initials containing animated scenes Ornamented initial (78,274 instances): Decorated initials with penwork, champie, or floral motifs Simple initial (23,121 instances): Colored initials without additional decoration Marginal medallion (878 instances): Framed marginal scenes Marginal decoration (1,624 instances): Unframed animated marginal elements (zoomorphic, heraldic) Annotation Methodology Annotations were created through an iterative human-in-the-loop process: Manual annotation of subset A by domain experts Training of detection models Automatic pre-annotation of new manuscripts Systematic manual validation with multiple quality control measures The annotation ontology was specifically designed to balance traditional codicological terminology with computational constraints and large-scale annotation requirements. All decisions are documented in the accompanying article. Data Sources Images were downloaded from IIIF-compliant digital libraries, primarily: Institut de Recherche et d'Histoire des Textes (IRHT-CNRS): ~12,230 images Bibliothèque nationale de France (Gallica): ~1,090 images The Walters Art Museum (via Stanford): ~236 images 30+ additional institutions across Europe and North America All images include metadata with original IIIF URLs, dimensions, manifest references, and canvas identifiers, ensuring full traceability and reproducibility. File Formats Images: JPEG format, minimum dimension 1200 pixels Annotations: YOLO format (TXT files with normalized coordinates) Metadata: CSV files with comprehensive image and annotation information Visualizations: HTML files for browsing annotations by class or manuscript Technical Specifications Annotation type: Rectangular bounding boxes (YOLO format) Coordinate system: Normalized (0-1) relative to image dimensions Resolution: Downloaded at maximum IIIF resolution, resized to min. 1200px File naming: Systematic pattern enabling manuscript identification Quality control: Multiple validation passes, overlap detection, visual inspection Use Cases This dataset enables: Digital Humanities Research: Quantitative codicology: systematic analysis of decoration density and distribution Comparative studies across manuscripts, regions, and time periods Statistical analysis of decorative programs Heritage Applications: Automated indexing of manuscript collections Enhanced search and discovery interfaces Digital exhibition curation Computer Vision Research: Object detection in historical documents Handling of class imbalance in cultural heritage data Transfer learning from medieval manuscripts Art History: Large-scale iconographic analysis Study of workshop practices and artistic attribution Analysis of text-image relationships Performance Benchmarks Models trained on this dataset achieve: Weighted F1 > 0.90 on majority classes (reflecting real-world annotation workflows) Standard F1: 0.64-0.80 across all classes (accounting for minority class challenges) Precision > 0.88 and Recall > 0.92 on ornamented initials See related publication (HORAE Detection Models) for detailed performance analysis. Related Publications HORAE_LSv2 Dataset (predecessor): https://doi.org/10.5281/zenodo.16919911 HORAE Detection Models (trained on this dataset): https://doi.org/10.5281/zenodo.17279775 Boillet et al. (2019): "HORAE: an annotated dataset of books of hours", HIP'19 @article{bernard2025horae_detection, title = {{Detection of Miniatures and Initials in Medieval Books of Hours: Ontology, Datasets, and Models HORAE_Minit}}, author = {Bernard-Leterme, Lise and Stutzmann, Dominique}, journal = {Humanités numériques}, year = 2025, note = {submitted}}



