遇见数据集

TranscriboQuest 2025: Medieval Latin

收藏
Zenodo2025-09-11 更新2026-05-26 收录
官方服务:

资源简介:

TranscriboQuest 2025: Medieval Latin Context of the dataset This dataset was created in the context of TranscriboQuest 2025 (Medieval Latin Team) held in Lyon (03/09/2025-05/09/2025). The goal of this summer school was to get more acquainted with eScriptorium and produce a limited, but qualitative dataset of medieval manuscripts. The aim of this dataset was to contribute to underrepresented aspects of the manuscripts used in the CATMuS project. We opted to focus on medieval glosses, in latin language. Documents The dataset comprises annotated images drawn from 12 manuscripts: a total of 37 folios, with 5060 segmented lines, including 735 transcribed lines. There is a total of 358 zones. The manuscripts include: Basel, Universitätsbibliothek, A VII 3. Cambridge, Trinity College, 153 (B.5.6). Florence, Biblioteca Medicea Laurenziana, Plut.78.16. Grenoble, Bibliothèque municipale, Ms.371 Rés. Oxford, Bodleian Library, Ms. Laud. Lat. 53. Paris, Bibliothèque nationale de France, Département des manuscrits, Latin 268. Paris, Bibliothèque nationale de France, Département des manuscrits, Latin 7900A. Paris, Bibliothèque nationale de France, Département des manuscrits, Latin 11560. St. Gallen, Kantonsbibliothek, Vadianische Sammlung, VadSlg Ms. 312 Strasbourg, Bibliothèque nationale et universitaire, Ms.0.044. Vatican, Biblioteca Apostolica Vaticana, Pal.lat.1580. Vatican, Biblioteca Apostolica Vaticana, Vat.lat.3083. Segmentation The team focused mainly on segmentation with a limited number of pages transcribed and corrected. Some team members chose to detect manuscript zones and segment lines with the blla.mlmodel while others with more difficult layout tagged zones manually before running that same model for the automatic detection of lines Baselines and Masks. All identified zones and lines were then manually refined. The greatest difficulty for the model lay in the diverse layouts of medieval manuscripts. It frequently merged several columns into a single text zone, which required manual correction. It also missed a number of interlinear additions and found non-existing lines in illuminations and decorated initials. Segmentation guidelines used are based upon SegmOnto, a controlled vocabulary to describe the layout of pages. The main rules we followed include: Zone tagging with "MainZone," "MarginTextZone," and in some cases "DamageZone," "DropCapitalZone," "NumberingZone," "RunningTitleZone," "QuireMarksZone," "StampZone," and "GraphicZone." "Numbering zone" is used to define the area with the page number. Line tagging with "DefaultLine," and "InterlinearLine." As for initials, in some manuscripts they are separated blocks, while in others they are considered as the other letters in the line.In some case, subtypes were used for zones or lines, for instance: "CustomZone:Title." Two different files are provided. Given `{name}` file:- `{name}.subclasses.xml` contains the annotations that keep the regions and lines classes and subclasses (e.g. `MarginTextZone` and `MarginTextZone:auctoritas_span`)- `{name}.classes.xml` reduces each subclass to its parent class. In the above example, all the marginal zones will be considered and annotated as `MarginTextZone`. Transcription The CATMuS guidelines were followed when transcribing the documents.To initially transcribe the manuscripts, CATMuS Medieval 1.6.0, Pinche et al., was applied.All errors were corrected manually and the provided transcriptions can be considered Ground Truth. Abbreviations were not expanded. Creators of the dataset Agnès Boutreux, Romain Chevalier, Chiara Corongiu, Sarah Gaucher, Estelle Guéville, Annabelle Kienzl, Jan Maliszewski under supervision of Matthias Gille Levenson.

提供机构:
Zenodo
创建时间:
2025-09-05
二维码
社区交流群
二维码
科研交流群
商业服务