遇见数据集

Portuguese Handwriting 16th-19th c.

收藏
Zenodo2025-02-20 更新2026-05-26 收录
官方服务:

资源简介:

All data were imported from the platform Transkribus on which the AI model for automatic transcription “Portuguese Handwriting 16th-19th c.” was last trained in July 2023 with the recognition engine Pylaia, and can now be used. The data are divided into ten folders, according to the total number of the trainings, from the initial to the definitive one, plus one set for final validation. The eight previous trainings were realized between June 2022 and May 2023. The history of all trainings can be read on e-Inquisition. Each of these folders corresponds to one collection in the platform; every collection has a number of documents; every document has a number of images, or pages, as indicated below. The ten uploaded folders (zip) are distributed as follows: —nine Training Sets (TS) (ca 92% of the whole data; status of the transcriptions from the TS: Ground Truth); —the final Validation Set (VS) (ca 8% of the whole data; status of the transcriptions from the VS: Ground Truth). All TS folders contain only the new data added to the following training (thus added to the previous data). Only the last VS, which is complete (505 p.), is provided. One document = images / transcribed pages (Ground Truth: transcription made by the members of TraPrInq project (Transcrever os processos da Inquisição portuguesa, 1536-1821 | Transcribing the court records of the Portuguese Inquisition, 1536-1821), which lasted from January 2023 to July 2024. The majority of the documents are titled as follows: IL_number = document extracted from a trial record (processo) by the Inquisition of Lisbon_number of the processo; other titles: IC_ = Inquisition of Coimbra; IE_ = Inquisition of Évora. Total of transcribed pages: 6,417. The quality of the images in the data (jpg) is equal to that of the images used for automatic transcription. All digitized images can be found on the catalog of the Portuguese National Archives (Arquivo Nacional da Torre do Tombo, ANTT). Available data (10 zip files, total size 6.7 GB): Training Set1: 698 pages/images Training Set2: 984 pages/images Training Set3: 869 pages/images Training Set4: 926 pages/images Training Set5: 631 pages/images Training Set6: 665 pages/images Training Set7: 564 pages/images Training Set8: 549 pages/images Training Set9: 531 pages/images Validation Set_Final: 505 pages/images 2-one pdf file: Paleographical criteria used by the team for the transcription of the documents; list of characters (in Portuguese).

所有数据均导入自Transkribus平台:该平台于2023年7月使用识别引擎Pylaia完成了自动转录模型“16至19世纪葡萄牙手写体(Portuguese Handwriting 16th-19th c.)”的最后一次训练,目前该模型已可投入使用。 本数据集依据训练总轮次分为十个文件夹,涵盖从初始训练至最终训练的全部轮次,另附带一套最终验证集。此前共完成8轮训练,时间跨度为2022年6月至2023年5月。所有训练的历史记录均可在e-Inquisition平台查阅。每个文件夹对应平台内的一个馆藏集,每个馆藏集包含若干文档,每份文档则包含若干图像(即页面),详情如下。 本次上传的10个压缩包(ZIP格式)分布如下: ——9个训练集(Training Sets, TS),占总数据量的约92%;训练集的转录标注状态为真实标签(Ground Truth); ——1个最终验证集(Validation Set, VS),占总数据量的约8%;验证集的转录标注状态为真实标签(Ground Truth)。 所有训练集文件夹仅包含对应训练轮次新增的数据(即基于此前已有数据补充的新样本)。 本次仅提供完整的最终验证集,共计505页。 每份文档对应若干图像/转录页面,其真实标签(Ground Truth)由TraPrInq项目团队制作。该项目全称为“转录葡萄牙宗教裁判所档案(1536-1821)(Transcrever os processos da Inquisição portuguesa, 1536-1821 | Transcribing the court records of the Portuguese Inquisition, 1536-1821)”,执行周期为2023年1月至2024年7月。 大部分文档的命名规则如下:IL_编号代表从里斯本宗教裁判所的审判档案(葡萄牙语:processo)中提取的文档,编号为档案编号;其余命名格式包括:IC_代表科英布拉宗教裁判所相关文档;IE_代表埃武拉宗教裁判所相关文档。 总转录页面数:6417页。 数据集内的JPEG格式图像质量与自动转录所用图像的质量一致。 所有数字化图像均可在葡萄牙国家档案馆(Arquivo Nacional da Torre do Tombo, ANTT)的馆藏目录中查阅。 本次可用数据集包含10个压缩包,总大小为6.7 GB,具体如下: 训练集1:698页/图像 训练集2:984页/图像 训练集3:869页/图像 训练集4:926页/图像 训练集5:631页/图像 训练集6:665页/图像 训练集7:564页/图像 训练集8:549页/图像 训练集9:531页/图像 最终验证集:505页/图像 附带1个PDF文件:包含团队用于文档转录的古文字学标注准则,以及葡萄牙语字符表。

提供机构:
Zenodo
创建时间:
2024-10-25
二维码
社区交流群
二维码
科研交流群
商业服务