Honkoku-Lines: A Million-Scale Line-Level Image–Transcription Dataset for Historical Japanese OCR Built from Crowdsourced Transcriptions
收藏资源简介:
This record archives the construction source code and the line-level companion metadata for Honkoku-Lines, a line-level dataset for recognition of handwritten historical Japanese (kuzushiji): images of individual text lines cut from premodern documents, each paired with a diplomatic transcription of that line. It is derived from Minna de Honkoku, a citizen-science transcription platform, by automatically aligning crowdsourced page-level transcriptions with line regions detected in the corresponding IIIF page images. The line images and the loadable dataset are published on Hugging Face: https://huggingface.co/datasets/yuta1984/honkoku-lines. This record provides the accompanying artefacts: src.zip — the construction pipeline as it was run, in execution order: IIIF page retrieval, RTMDet line detection, dual-margin cropping, PARSeq recognition, edit-distance alignment against the crowdsourced transcriptions, WebDataset packing, and the release and quality-review tooling (MIT). lines.jsonl.gz — per-line metadata for all 1,169,304 lines, images excluded: item, page and line identifiers, bounding box, IIIF region URL, transcription with its original notation, notation-stripped text, reference recogniser output, alignment distance and character-count ratio, detector confidence, image license, holding institution, and the recommended train/val/test split. items.tsv — per-item bibliography, counts and license (4,140 items). hosts.tsv — per-provider summary over the 31 IIIF endpoints. koji_notation.md — specification of the transcription notation. The dataset covers 1,169,304 lines from 4,140 items and 79,086 pages, 18.2 million transcribed characters, held by 29 institutions and aggregators. Line images for 1,018,511 of those lines (87.1 %) come from items whose rights permit redistribution and are bundled on Hugging Face; for the remainder, the IIIF region URLs in lines.jsonl.gz allow the images to be fetched from the holding institutions. Transcriptions and metadata are released under CC BY-SA 4.0 and the code under MIT; line images carry the license of the holding institution, stated per record. See README.md in this record for the full description, schema and limitations.



