遇见数据集

NewsEye / READ OCR training dataset from French Newspapers (18th, 19th, early 20th C.)

收藏
Zenodo2021-03-12 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The dataset comprises French newspaper pages from 18th, 19th and early 20th century with carefully corrected text. The page images were provided by the French National Library and comprise 127 pages (training set) and 8 pages (validation set). The data are formed according to the PAGE format (cf. Cf. https://github.com/PRImA-Research-Lab/PAGE-XML/) and were produced with the Transkribus platform with support of the NewsEye and the READ project.

本数据集收录18世纪、19世纪及20世纪早期的法国报纸版面,所有文本均经过精心校正。该数据集的版面图像由法国国家图书馆提供,其中训练集包含127个版面,验证集包含8个版面。本数据集采用PAGE格式(PAGE format,参见https://github.com/PRImA-Research-Lab/PAGE-XML/)进行构建,通过Transkribus平台制作,并得到了NewsEye与READ项目的支持。

提供机构:
Zenodo
创建时间:
2020-11-27
二维码
社区交流群
二维码
科研交流群
商业服务