遇见数据集

Turkish OCR Text Image Dataset: Nutuk, Wikipedia, and Gold Standard

收藏
Zenodo2026-08-14 更新2026-08-20 收录
官方服务:

资源简介:

This dataset contains three subsets of Turkish text images for OCR training and evaluation:Nutuk (Synthetic) – 6600 synthetic Turkish text images generated from Mustafa Kemal Atatürk's Nutuk.Wikipedia (Synthetic) – 6600 Synthetic Turkish text images generated from Wikipedia articles.Gold Standard (Real) – 90 Real-world Turkish text images collected under natural conditions.This dataset is openly available for academic and research purposes under the CC BY 4.0 license.

提供机构:
Zenodo
创建时间:
2026-08-13
二维码
社区交流群
二维码
科研交流群
商业服务