遇见数据集

HTR Winter School 2025 - Medieval Czech - Krakow, Biblioteka Jagiellońska, shelfmark BJ Rkp. 441 IV

收藏
Zenodo2026-01-08 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains ground truth based on the the Kronika česká do roku 1330 (1531, Krakow, Biblioteka Jagiellońska, shelfmark BJ Rkp. 441 IV, available from: https://jbc.bj.uj.edu.pl/dlibra/publication/361505/edition/344903#description, Old Czech) and it was prepared by participants of the Czech group during Winter School: HTR of Historical Sources, held in Vienna in 2025. Manuscript: Old Czech, Gothic Bastarda, beginning of the 16th C. Origin of the data Source of images: Jagiellońska Biblioteka Cyfrowa https://jbc.bj.uj.edu.pl/dlibra/publication/361505/edition/344903#description Transcription Rules for Pulkava's Chronicle: Transliteration General Principle: The text is transliterated in the edition, i.e., reproduced without changing the orthographic system. Limitations: The used "measure" of transliteration has its limits; it cannot be considered a complete transcription of all literal elements of the source document. Letter Forms: We do not distinguish different graphic forms of letters (simple `<v>`, `<v>` with a loop, etc.). We do distinguish the round `<s>` and the long `<ſ>`. We do not distinguish the dotted i `<i>` and the dotless i `<ı>`; these allographs are transcribed according to the current usage as a single letter `<i>`. We do not transcribe initials (Transkribus cannot recognize them and place them on the same line with the rest of the text). Punctuation & Diacritics: We keep the punctuation as it appears in the manuscript. We standardize various marks above the graphemes to a dot, caron (háček), or accent (čárka). We transcribe the caron according to the print, i.e., even `m̌`, `ň`, although in modern orthography we would write `ě`. We do not reconstruct vowel quantity. Capitalization: We write uppercase and lowercase letters according to the manuscript. Word Segmentation: In the transliterated transcription, we separate prepositions, conjunctions, and names. We do not separate emphatic particles `-ť`, `-ž`. Morphology: In word forms, we respect fluctuation and the simplification of geminates. Methodology: Unlike paleographic transcription, this is not an imitation of the manuscript (e.g., distinguishing various graphic forms of letters with the exception of distinguishing the long and round s), but a character-for-character transcription. Foliation: We do not transcribe foliation. Abbreviations: We tag the abbreviation with "abbreviation" and write it out in "expansion". We do not expand scribal abbreviations in parentheses according to usage, but transcribe them according to their form in the manuscript (e.g., `ḡt`); the loop is transcribed on the first letter as a horizontal line. If the abbreviation is in an inflectional ending, the last letters are written small above the baseline; we transcribe them in a superscript (e.g., `mnohéo`). Data organisation Selection of semi-diplomatic transcribed texts from the so-called Krakow, Biblioteka Jagiellońska, shelfmark BJ Rkp. 441 IV (1531). Texts were transcribed by the participants of the HTR Winter School 2025 in Vienna: Daniel Katscher: 2r; 14r Tereza Hejdová: 2v; 14v Veronika Jakubcová: 3r; 15r Marie Urbánková: 3v; 15v Jan Dušek: 4v; 16r Barbora Doušková: 5v; 16v Karolina Švecová: 6r; 17r Veronika Nová: 6v; 17v Linda Rudenka: 7r Martina Spěváčková: 7v; 18r Jana Mešková: 8r; 18v Markéta Pytlíková: 8v; 19r Laura Landová: 9r; 19v Michaela Reimannová: 9v; 20r Kateřina Kozlovská: 10r; 20v Vika Veličkaite: 10v; 21r Andrej Kostelník: 11r; 21v Anna Javoříková: 12r; 22r Diana Kostelníková: 12v; 23r Věra Soukupová: 13v; 23v and corrected by Anna Michalcová, Jitka Filipová and Eliška Pěnkavová. Number of transcribed pages is listed above. How to cite This dataset was created by Barbora Doušková, Jan Dušek, Jitka Filipová, Marie Hedvíková, Tereza Hejdová, Veronika Jakubcová, Anna Javoříková, Daniel Katscher Andrej Kostelník, Diana Kostelníková, Kateřina Kozlovská, Laura Landová, Jana Mešková, Anna Michalcová, Veronika Nová, Eliška Pěnkavová, Markéta Pytlíková, Michal Racyn, Michaela Reimannová, Věra Soukupová, Martina Spěváčková, Karolina Švecová, Marie Urbánková, Vika Veličkaite. The digitisation is not copyright free, but the transcription is. However, properly annotating a corpus takes time and is a task that should be recognised. If you use any item from this corpus as ground truth, cite the dataset using the following information Copyright and licence This dataset was created as part of the Winter School of HTR of Historical Sources 2025, Vienna at the Österreichische Akademie der Wissenschaften, Institut für Mittelalterforschung, all transcriptions are licensed under the Creative Commons 4 licence. Images were provided by the Austrian National Library (ÖNB) and are licensed under Creative Commons 4 licence.

提供机构:
Zenodo
创建时间:
2026-01-08
二维码
社区交流群
二维码
科研交流群
商业服务