遇见数据集

Dataset for BnF, fr. 2813 - Grandes Chroniques de France

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the extended version of the ground truth for the codex Paris, BnF, fr. 2813, used in the experiments for the paper “Leveraging Morphology for Metrological Historical Script Analysis”, accepted to International Conference on Document Analysis and Recognition (ICDAR 2026, Vienna, Austria). What’s New Compared to v.1 95 newly annotated folios have been added (see the new btv1b84472995_metadata.csv for details); The ALTO XML annotations now distinguish between #MainZone#1 and #MainZone#2, corresponding to the column order on each page; Two versions of annotation.json are provided: one version includes hyphenation for word breaks at the end of lines. As for version 1, the repository is organized into two main data folders: --- 📁 `btv1b84472995_GT.zip` This folder contains the ground truth dataset used for Handwritten Text Recognition (HTR), created from the selected folia of the manuscript Paris, BnF, français. 2813 The identifier `btv1b84472995` refers to the ark ID of this manuscript in Gallica. Folder structure: btv1b84472995_GT ├── images └── annotations - `images/`: High-resolution selected images downloaded from Gallica. Image names follow the pattern `btv1b84472995_f<number>`, corresponding to the Gallica view number. ➤ Credit: *Source gallica.bnf.fr / Bibliothèque nationale de France* - `annotations/`: XML-ALTO annotation files created with eScriptorium. Layout: Annotations follow the Segmonto ontology. The potential users of the ground truth should note that we use additional personalized tags for: - `'RubricLines'`: Rubricated lines - `'HalfLines'`: Partial or incomplete lines - `'MainZone#1'` and `'MainZone#2'`: order of the column, instead of simply #MainZone Transcription: The dataset is CATMuS-compliant, using a graphemic transcription approach. --- 📁 `dataset.zip` This folder contains the dataset used in the experiments described in the paper, using the DTLR architecture for paleography, as detailed in the paper. Folder structure: dataset ├── images └── annotation.json - `images/`: Each subfolder contains polygonal line extractions (with alpha transparency) per manuscript page.- `annotation.json`: Contains the annotation and metadata for each line. `annotation.json` structure example: ```json"<image_id>": { // corresponds to the image names in the images folders "split": "train", "label": "A beautiful calico cat.",// Transcription text of the line "line": "DefaultLine", // Type of line "zone": "MainZone#1", // Type of Zone where the line is found "script": "RaouletOrleans", // Identifier for the scribal hand "folio": "1r", "gp": "GP1", // Identified Graphic Profile "doc": "HT1", } Papers associated with the data: v1: https://malamatenia.github.io/bnf-fr-2813/ (Scriptorium 2026) v2: https://malamatenia.github.io/dtlr-for-metrology/ (ICDAR 2026)This study was supported by the CNRS through MITI and the 80|Prime program (CrEMe Caractérisation des écritures médiévales), and by the European Research Council (ERC project DISCOVER, number 101076028).

本仓库包含《Leveraging Morphology for Metrological Historical Script Analysis》论文实验中所用的巴黎国家图书馆藏抄本BnF, fr. 2813的基准真值(ground truth)扩展版本,该论文已被国际文档分析与识别会议(International Conference on Document Analysis and Recognition, ICDAR 2026,奥地利维也纳)收录。 与v1版本相比的更新内容 - 新增95张带标注的抄本页叶,详情请见新增的`btv1b84472995_metadata.csv`文件; - ALTO XML标注现已区分`#MainZone#1`与`#MainZone#2`,分别对应每页的两栏排版顺序; - 本次提供两个版本的`annotation.json`:其中一个版本包含行尾单词断连的连字符标注。 与v1版本一致,本仓库分为两个主要数据文件夹: --- 📁 `btv1b84472995_GT.zip` 本文件夹包含手写文本识别(Handwritten Text Recognition, HTR)所用的基准真值数据集,源自巴黎国家图书馆藏抄本fr. 2813的精选页叶。标识符`btv1b84472995`为该抄本在Gallica平台的ARK编号。 文件夹结构: btv1b84472995_GT ├── images └── annotations - `images/`:从Gallica平台下载的高分辨率精选图像。图像命名格式为`btv1b84472995_f<编号>`,对应Gallica平台的视图编号。 ➤ 版权来源:*Source gallica.bnf.fr / Bibliothèque nationale de France* - `annotations/`:使用eScriptorium工具生成的XML-ALTO格式标注文件。 标注布局:遵循Segmonto本体规范。基准真值的潜在使用者需注意,本项目使用了以下自定义标签: - `'RubricLines'`:红字装饰标题行 - `'HalfLines'`:半行或不完整行 - `'MainZone#1'` 与 `'MainZone#2'`:表示栏目的排版顺序,而非仅使用`#MainZone`单一标签 转录规范:本数据集符合CATMuS标准,采用字素级转录方案。 --- 📁 `dataset.zip` 本文件夹包含论文所述实验所用的数据集,采用面向古文字学的DTLR架构,详情请见论文。 文件夹结构: dataset ├── images └── annotation.json - `images/`:每个子文件夹包含对应手抄本页叶的多边形线条提取图像(带Alpha透明通道)。 - `annotation.json`:包含每一行文本的标注与元数据。 `annotation.json`结构示例: json { "<image_id>": { // 对应images文件夹中的图像文件名 "split": "train", // 数据集拆分类型 "label": "A beautiful calico cat.",// 该行文本转录内容 "line": "DefaultLine", // 行类型 "zone": "MainZone#1", // 该行所在的区域类型 "script": "RaouletOrleans", // 抄写人手迹标识 "folio": "1r", // 抄本页叶编号 "gp": "GP1", // 已识别的图形特征轮廓(Graphic Profile) "doc": "HT1" // 文档标识 } } 本数据集相关论文: - v1版本:https://malamatenia.github.io/bnf-fr-2813/ (Scriptorium 2026) - v2版本:https://malamatenia.github.io/dtlr-for-metrology/ (ICDAR 2026) 本研究获法国国家科学研究中心(CNRS)通过MITI与80|Prime项目(CrEMe 中世纪手写文字特征分析),以及欧洲研究理事会(ERC项目DISCOVER,编号101076028)资助。

提供机构:
Zenodo
创建时间:
2025-04-28
二维码
社区交流群
二维码
科研交流群
商业服务