nagohachi/NDL_pdm-ocr-part2_cropped
收藏资源简介:
该数据集是日本国立国会图书馆发布的公共领域OCR训练数据集(FY2021)的衍生版本。每个样本是历史文档页面的行级裁剪图像,与其OCR文本转录配对。数据集经过变换:从原始页面图像中根据边界框坐标裁剪出每个文本行,并将裁剪图像与行的转录文本配对,所有样本打包成WebDataset的.tar分片。数据集包含103,882个样本,分为4个分片,覆盖时期为1870年代至1960年代,源图像来自3,997个XML文件/7,919个页面图像,样本键格式为{decade}/{book_id}/{page_name}/{line_index}。每个样本包含JPEG格式的裁剪行图像(质量95)和文本转录文件。
This dataset is a derivative of the Public Domain OCR Training Dataset (FY2021) published by the National Diet Library of Japan. Each sample is a line-level crop of a historical document page, paired with its OCR text transcription. The dataset was transformed by cropping each text line from the original page image using its bounding box coordinates and pairing the cropped image with the lines transcription text, with all samples packaged into WebDataset .tar shards. It contains 103,882 samples across 4 shards, covering the period from the 1870s to the 1960s, sourced from 3,997 XML files / 7,919 page images, with sample key format {decade}/{book_id}/{page_name}/{line_index}. Each sample includes a cropped line image in JPEG format (quality 95) and a text transcription file.



