Standardized Labels and Splits of OCR Datasets
收藏官方服务:
资源简介:
Standardized labels and data splits for various OCR datasets to facilitate consistent evaluation and benchmarking of OCR models.This dataset includes harmonized annotations and predefined training, validation, and test splits for multiple OCR datasets. Included datasets: - IAM-line (handwriting) - TotalText (curved scene text) - TextOCR (scene text) - SynthText (synthetic overlays) - TE-synthetic (medical synthetic text) - MultiCaRe (medical overlaid captions) - PHI (imprints including protected health information) - SROIE (receipts for KIE) - POIE (product labels for KIE)
提供机构:
Zenodo创建时间:
2025-11-27



