Sindhi OCR Dataset: Multi-Font Machine-Printed Text Line Images with Ground Truth for Low-Resource Language Recognition
收藏官方服务:
资源简介:
This dataset has 60,122 binary TIF text line images across different Sindhi fonts. Ground truth provided as UTF-8 encoded .txt files and .xml metadata files. Train/test splits provided as split files for reproducibility. For academic and non-commercial research use only.
提供机构:
Zenodo创建时间:
2026-05-18



