Groundtruth data for HTR from the missionary Chinese-Latin dictionary Rinuccini 22
收藏资源简介:
Dataset description This Automatic Text recognition (ATR) groundtruth data is from the Chinese-Latin dictionary Rinuccini 22, held in the Medicea Laurenziana Library in Florence and is published with the permission of the library. This dictionary, compiled by Basilio Brollo da Gemona (1648–1704), is organized by the radicals and strokes of Chinese characters. As part of the research project ChEDiL (Projet-ANR-23-CE27-0008), the Latin-script portion (romanization and translations) has been fully transcribed and/or verified on e-Scriptorium by Michiel Streijger (University of Regensburg). He received the support of Shad Mohammad (supervised by Elod Egyed-Zsigmond, INSA-Lyon) to achieve this task. Chen Fengyi (EFEO / ChEDiL), supervised by Vincent Paillusson (HTL-UMR 7597), prepared the dataset. It is composed of 300 transcribed pages (9434 lines) in Latin script to transcribe both Latin language and Chinese pronunciation using Latin script combined with many, not standardized at the time, diacritical marks. The transcription of the introductory texts from this volume and two other related dictionaries will be deposited on HAL by Michela Bussotti (EFEO UMR CCJ), Cécile Tep (EFEO/University of Caen Normandy), and Michiel Streijger. Data organization The organization of the dataset is pretty straight forward: the "image" directory contains all the line-image files extracted and binarized from the dictionary. Their names are randomized. the "label_line" directory contains files with the textual content of each image expressed in a json format.



