Synthetic OCR Dataset: 105,738 Tamil Text Lines Rendered in 27 Diverse Fonts with Corresponding Ground Truth Annotations
收藏资源简介:
This dataset provides a benchmark for Tamil Optical Character Recognition (OCR), covering both handwritten (Hangual) and printed Tamil text. It includes high-quality ground truth (GT) text files paired with corresponding TIFF images, making it valuable for training and evaluating OCR models, particularly for Tesseract, deep learning-based recognition, and AI research. Dataset Highlights Total Size: 3.6GB Total Pairs: Approximately, 7,69,396 tiff, gt.txt and box files Handwritten Fonts (9 Unicode Fonts): Aazhi, Gnani, Hemalatha, Indumathi, Kalayarasi, Siva_01, Siva_02, Sudeeptha, Yogeshwaran Printed Fonts (9 Unicode Fonts): Arima, HindMadurai, MuktaMalar, TAU-Ezhil, TAU-Kambar, TAU-Marutham, TAU-Neythal, TAU-Valluvar, AnekTamil, Catamaran, Kavivanar, Pavanam, TAU-Barathi, TAU-Kabilar, TAU-Malar, TAU-Mullai, TAU-Nilavu, TiroTamil Data Source: The text corpus (GT text files) is curated from Wikipedia, Wikisource, Maattru and Theekkathir ensuring linguistic diversity. The fonts are publicly available Unicode Tamil fonts, sourced from Google Fonts and Tamil Virtual University. File Structure tam_new-ground-truth/├── 00001.gt.txt├── 00001.tiff├── 00001.box├── 00001.lstm├── 00002.gt.txt├── 00002.tiff├── 00002.box├── 00002.lstm├── ... Cite this work @dataset{tamilocr_dataset_2025, author = {Syedkhaleel Jageer}, title = {Synthetic OCR Dataset: 105,738 Tamil Text Lines Rendered in 27 Diverse Fonts with Corresponding Ground Truth Annotations}, year = {2025}, publisher = {Zenodo}, doi = {10.5281/zenodo.16881612}, url = {https://doi.org/10.5281/zenodo.16881612}}



