遇见数据集

Sindhi OCR Dataset: Multi-Font Machine-Printed Text Line Images with Ground Truth for Low-Resource Language Recognition

收藏
Zenodo2026-05-18 更新2026-05-26 收录
官方服务:

资源简介:

This dataset has 60,122 binary TIF text line images across different Sindhi fonts. Ground truth provided as UTF-8 encoded .txt files and .xml metadata files. Train/test splits provided as split files for reproducibility. For academic and non-commercial research use only.

提供机构:
Zenodo
创建时间:
2026-05-18
二维码
社区交流群
二维码
科研交流群
商业服务