Ground Truth Training Dataset for Tibetan Generic 0.1 HTR Model for Tibetan Newspapers
收藏资源简介:
This ground truth dataset comprises the training data for Tibetan Generic 0.1, the first version of a generic Handwritten Text Recognition (HTR) model for modern Tibetan. The model is designed to recognize Tibetan Uchan (dbu can) and Ume (dbu med) scripts, as well as limited English and Chinese text. The corpus spans materials from the 18th to the 20th century, drawing from diverse sources: legal documents (contributed by Daniel Wojahn), modern books published between the 1950s and 1980s (Divergent Discourses project), and Tibetan-language newspapers from the 1950s and 1960s (Divergent Discourses project). All training data were manually annotated in Transkribus by human transcribers working within the Divergent Discourses project, which is jointly funded by the Arts and Humanities Research Council (AHRC, UK) and the Deutsche Forschungsgemeinschaft (DFG, Germany). A second annotator verified all transcriptions to ensure quality. Dataset Statistics: 161,482 words 1,533 pages The resulting HTR model, Tibetan Generic 0.1 (Transkribus ID 373545), achieves a Character Error Rate (CER) of 3.58% and is publicly available on the Transkribus platform. The dataset includes original page images alongside transcriptions in PAGE XML format. Important Note: This dataset has been verified for HTR training purposes only. Annotations of text regions, baselines, and line polygons have not been systematically reviewed and should be used accordingly.



