遇见数据集

Page Extraction Dataset

收藏
Zenodo2026-01-22 更新2026-05-26 收录
官方服务:

资源简介:

In digitised cultural heritage items such as books, newspapers and archival records, a problem that can negatively affect OCR are black margins around a page caused by document scanning. In order to enable document layout analysis (DLA), these black margins need to be cropped and the pages need to be extracted correctly. To enable the training of a machine learning model capable of extracting pages, a dataset was created. The machine learning task for which this dataset was collected falls into the domain of image segmentation and, more generally, of computer vision. The dataset was compiled by Vahid Rezanezhad within the research project “Mensch.Maschine.Kultur – Künstliche Intelligenz für das Digitale Kulturelle Erbe” at the Staatsbibliothek zu Berlin – Berlin State Library (SBB). The research project was funded by the Federal Government Commissioner for Culture and the Media (BKM), project grant no. 2522DIG002. The Minister of State for Culture and the Media is part of the German Federal Government.

创建时间:
2026-01-22
二维码
社区交流群
二维码
科研交流群
商业服务