遇见数据集

Synthetic Arabic Gazette-Style Text-Line Images for OCR Training (315k lines)

收藏
Zenodo2026-07-31 更新2026-08-02 收录
官方服务:

资源简介:

A corpus of 315,000 synthetic Arabic text-line images in the style of officialgazette print (300,000 train / 5,000 validation / 10,000 test), each pairedwith the transcription it was rendered from (UTF-8 JSONL). Lines are renderedwith fonts, layout statistics, and degradation levels matched to real gazettescans, providing large-scale training data for Arabic printed-text linerecognition under realistic degradation. The record includes the fixed 59-character inventory used for rendering(char_vocab.json) and SHA-256 checksums. Some renders are clipped at thecanvas edge (a density heuristic flags 25.4% of a training sample as an upperbound; see README.md); the corpus is published exactly as generated and used,without post-hoc filtering, and the README gives a simple filtering recipe forconsumers who prefer a conservative subset. Files: images_train_*.zip (300k PNG line crops, sharded), images_val.zip (5k),images_test.zip (10k), {train,val,test}_rendered.jsonl, char_vocab.json,README.md, checksums.txt. License: CC BY 4.0. Fully synthetic renders; no scanned or personal material.

提供机构:
Zenodo
创建时间:
2026-07-31
二维码
社区交流群
二维码
科研交流群
商业服务