sinhala_synthetic_ocr_news_large
收藏资源简介:
Sinhalaa数据集是一个包含僧伽罗语(Sinhala)文本和对应图像的大规模合成数据集,专门为光学字符识别(OCR)任务设计。数据集来源于Kaggle平台上的“sinhala-ocr-image-creation”项目,通过合成方法生成。数据集中包含80,000个训练样本,每个样本由两个字段构成:image字段存储图像数据,text字段存储对应的僧伽罗语文本字符串。数据集总大小约为16.5GB,仅提供训练分割。该数据集适用于僧伽罗语OCR模型的训练、评估与研究,旨在提升对僧伽罗语文本的图像识别能力。
The Sinhalaa Dataset is a large-scale synthetic dataset containing Sinhala text and corresponding images, specifically designed for optical character recognition (OCR) tasks. It is derived from the 'sinhala-ocr-image-creation' project on the Kaggle platform and generated via synthetic methods. The dataset includes 80,000 training samples, each composed of two fields: the 'image' field stores image data, and the 'text' field stores the corresponding Sinhala text string. The total size of the dataset is approximately 16.5 GB, and only the training split is provided. This dataset is suitable for the training, evaluation and research of Sinhala OCR models, aiming to improve the image recognition capability for Sinhala text.
数据集概述
- 数据集名称:sinhala_synthetic_ocr_news_large
- 语言:僧伽罗语(si)
- 数据集大小:约 16.48 GB(下载大小约 16.49 GB)
- 配置:仅包含
default配置,训练数据文件位于data/train-*
数据特征
- image:图像数据,类型为
image - text:文本数据,类型为
string
数据划分
- 训练集(train):共 80,000 个样本,占用约 16.48 GB 存储空间
来源与引用
- 创建来源:基于 Kaggle 代码 sinhala-ocr-image-creation 生成
- 引用格式:提供 BibTeX 引用,作者为 Ransaka R.(2026 年),DOI 为 10.57967/hf/9748




