遇见数据集

GaborMadarasz/synthetic_OCR_dataset_v2

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

匈牙利合成OCR数据集是一个用于微调匈牙利语文本光学字符识别(OCR)模型的合成生成图像-文本数据集。该数据集特别注重正确识别匈牙利语特有的重音字符(如á、é、í、ó、ö、ő、ú、ü、ű)。真实文本来源于匈牙利维基百科,图像通过PIL和Albumentations程序化渲染生成,并模拟了四种视觉配置文件以覆盖真实世界文档条件:现代(modern)、扫描(scanned)、档案(archiv)和段落(paragraph)。数据集包含约100,000张图像,采用CC BY-SA 4.0许可证,格式为HuggingFace数据集,包含图像特征和JSON对话格式。

The Hungarian Synthetic OCR Dataset is a synthetically generated image–text dataset for fine-tuning OCR models on Hungarian text, with a particular focus on correct recognition of Hungarian-specific accented characters (á, é, í, ó, ö, ő, ú, ü, ű). The ground-truth text is sourced from Hungarian Wikipedia. Images are rendered programmatically using PIL and augmented with Albumentations, across four visual profiles that simulate real-world document conditions: modern, scanned, archiv, and paragraph. The dataset contains approximately 100,000 images, is licensed under CC BY-SA 4.0, and is formatted as a HuggingFace dataset with image features and JSON conversation format.

提供机构:
GaborMadarasz
二维码
社区交流群
二维码
科研交流群
商业服务