遇见数据集

ENC-PSL/MEDUSA_Synthetic_Lines

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含50万张合成行图像,为MEDUSA训练课程的Silver阶段生成。它由École nationale des chartes – PSL作为MEDUSA项目的一部分制作,用于多语言中世纪手写文本识别(HTR)。数据集旨在通过将文本资源渲染为逼真的合成手稿行图像,为MEDUSA模型提供语言和脚本级别的先验知识,特别是针对那些在现有图像-文本HTR语料库中代表性不足的中世纪语言(如日耳曼语、凯尔特语和斯拉夫语)。生成管道包括片段采样、渲染、墨水模拟、背景合成、扫描退化和文档效果等六个阶段,输出为JPEG图像和对应的JSONL元数据。数据集结构按语言组织,包含JPEG/ALTO XML文件对,适用于DocWorkflow工具。

This dataset contains 500,000 synthetic line images generated for the Silver stage of the MEDUSA training curriculum. It was produced at the École nationale des chartes – PSL as part of the MEDUSA project for multilingual medieval handwritten text recognition (HTR). The dataset operationalises text-only resources by rendering them as plausible synthetic manuscript line images, providing the MEDUSA model with lexical and script-level priors for languages that are underrepresented in existing image–text HTR corpora, particularly Germanic, Celtic, and Slavic medieval languages. The generation pipeline involves six stages: snippet sampling, rendering, ink simulation, background composition, scan degradation, and document effects, outputting JPEG images with corresponding JSONL metadata. The dataset is structured by language with paired JPEG/ALTO XML files, ready for use with DocWorkflow.

提供机构:
ENC-PSL
二维码
社区交流群
二维码
科研交流群
商业服务