Synthetic Arabic Gazette-Style Text-Line Images for OCR Training (315k lines)
收藏资源简介:
A corpus of 315,000 synthetic Arabic text-line images in the style of officialgazette print (300,000 train / 5,000 validation / 10,000 test), each pairedwith the transcription it was rendered from (UTF-8 JSONL). Lines are renderedwith fonts, layout statistics, and degradation levels matched to real gazettescans, providing large-scale training data for Arabic printed-text linerecognition under realistic degradation. The record includes the fixed 59-character inventory used for rendering(char_vocab.json) and SHA-256 checksums. Some renders are clipped at thecanvas edge (a density heuristic flags 25.4% of a training sample as an upperbound; see README.md); the corpus is published exactly as generated and used,without post-hoc filtering, and the README gives a simple filtering recipe forconsumers who prefer a conservative subset. Files: images_train_*.zip (300k PNG line crops, sharded), images_val.zip (5k),images_test.zip (10k), {train,val,test}_rendered.jsonl, char_vocab.json,README.md, checksums.txt. License: CC BY 4.0. Fully synthetic renders; no scanned or personal material.



