遇见数据集

Ground Truth for Medieval Greek Handwritten Text Recognition and Text Line Detection: Vat. gr. 2228 and Phil. gr. 130

收藏
Zenodo2026-06-16 更新2026-05-26 收录
官方服务:

资源简介:

Overview This repository provides PAGE-XML ground truth for two medieval Greek handwritten texts from the 14th-century manuscripts Vat. gr. 2228 and Phil. gr. 130. The files are intended for use in text line detection (TLD) and handwritten text recognition (HTR) workflows. Each PAGE-XML file contains text region and line geometry, including baselines and line polygons. For a subset of pages, the files also include line-level transcriptions. Because image rights remain with the holding institutions, manuscript images are not distributed through this repository. The corresponding image collections can be accessed at: Vat. gr. 2228 (Biblioteca Apostolica Vaticana): https://digi.vatlib.it/view/MSS_Vat.gr.2228.pt.1 and https://digi.vatlib.it/view/MSS_Vat.gr.2228.pt.2. Phil. gr. 130 (Österreichische Nationalbibliothek): https://viewer.onb.ac.at/13228923 For Phil. gr. 130, high-resolution images are freely available and correspond one-to-one with the PAGE-XML files in this release. The annotated images are .jpg files with dimensions ranging from 4032 to 4179 px in width and from 6040 to 6086 px in height. For Vat. gr. 2228, the publicly accessible images are low-resolution color reproductions with watermarking, whereas this PAGE-XML subset was prepared using higher-resolution grayscale images obtained directly from the library. The annotated source images used for the PAGE annotations are .tif files with dimensions of 2174 × 2996 px. The freely available online images are effectively downscaled versions of those source images, so line polygons and baselines must be scaled appropriately to align correctly. Transcription normalization The PAGE-XML transcriptions in this dataset have been lightly normalized to maintain the original transcriptions as closely as possible. In particular, the following normalization policy has been applied: remove zero-width and other invisible characters map Unicode whitespace to ASCII space (U+0020) strip leading and trailing spaces collapse consecutive ASCII spaces NB: For additional pre-processing and downstream tasks, NFC normalization is recommended to ensure a consistent Unicode representation. Coverage This release contains 46 PAGE-XML files in total: Phil. gr. 130: folios 86r-94v; 18 PAGE-XML files, all 18 with transcriptions. Vat. gr. 2228: folios 12r-15v, 17r-20v, and 194r-199v; 28 PAGE-XML files, 18 with transcriptions and 10 line-geometry-only. Dataset directory structure The dataset is organized into PAGE-XML files and split-definition CSV files as follows: pagexml/├── phil_gr_130/└── vat_gr_2228/splits/├── train_htr.csv├── train_line.csv├── val_shared.csv└── test_htr.csv The pagexml/ directory contains PAGE-XML annotation files grouped by manuscript. The splits/ directory contains CSV files defining dataset split membership. The split files are as follows: train_htr.csv: Listing the HTR training pages train_line.csv: Listing the TLD training pages val_shared.csv: Listing the validation pages shared by HTR and TLD test_htr.csv: Listing the HTR test pages In this context, the CSV files contain the following columns: manuscript: manuscript identifier (phil_gr_130 or vat_gr_2228) image_filename: image filename referenced by the PAGE-XML file xml_filename: PAGE-XML filename manuscript_folio: folio identifier

提供机构:
Zenodo
创建时间:
2026-04-05
二维码
社区交流群
二维码
科研交流群
商业服务