Turkish OCR Text Image Dataset: Nutuk, Wikipedia, and Gold Standard
收藏官方服务:
资源简介:
This dataset contains three subsets of Turkish text images for OCR training and evaluation:Nutuk (Synthetic) – 6600 synthetic Turkish text images generated from Mustafa Kemal Atatürk's Nutuk.Wikipedia (Synthetic) – 6600 Synthetic Turkish text images generated from Wikipedia articles.Gold Standard (Real) – 90 Real-world Turkish text images collected under natural conditions.This dataset is openly available for academic and research purposes under the CC BY 4.0 license.
提供机构:
Zenodo创建时间:
2026-08-13



