digital-history-bielefeld/image-text_anglicana-legal-texts
收藏资源简介:
该数据集包含21,579行高质量的手写文本识别(HTR)地面真实数据,由比勒费尔德大学FLOW项目(数字历史)创建,用于训练中世纪行政和法律文件的HTR模型。数据集包括图像片段(行)及其对应的转录和行坐标。材料聚焦于13世纪和14世纪用Anglicana字体和拉丁语书写的法律记录。每个条目包含:图像片段、行的外交转录、几何元数据(如行坐标和基线)、文档结构信息(如边注与正文)以及来源信息(如文件名和项目名称)。数字图像最初由AALT协会(可用租约和标题档案)提供,原始物理记录由英国国家档案馆(TNA)保存。转录根据Open Government Licence v3.0发布,图像片段经TNA特别许可发布,用于促进HTR研究和训练。数据集遵循外交转录方法,包括缩写扩展、字符规范化、原始拼写保留、特殊符号转录等。
This dataset contains 21,579 lines of high-quality Ground Truth data for Handwritten Text Recognition (HTR). It was created by the FLOW Project at Bielefeld University (Digital History) to facilitate the training of HTR models for medieval administrative and legal documents. The dataset consists of image snippets (lines) paired with their corresponding transcriptions and the line coordinates. The material focuses on legal records from the 13th and 14th centuries written in Anglicana script and Latin language. Each entry includes: the image snippet of a single line, the diplomatic transcription of the line, geometric metadata for precise localization (line_coords & line_baseline), information about the document structure (region_type), and provenance information (filename & project_name). The digital images were originally provided by the AALT Society (Archive of Available Leases and Titles), with original physical records held by The National Archives (TNA), Kew, UK. Transcriptions are published under the Open Government Licence v3.0, and image snippets are reproduced by permission of TNA for HTR research and training. The transcriptions follow a diplomatic approach, including expanded abbreviations, preserved original spelling, and specific rules for special signs.





