synthetic_OCR_dataset
收藏资源简介:
该数据集是一个合成生成的匈牙利文档OCR数据集,包含3000个匈牙利文档图像及其对应的真实文本转录。它专为训练和评估光学字符识别(OCR)、文档布局分析以及匈牙利行政和商业文档上的多模态语言模型而设计。数据集语言为匈牙利语,完全支持变音符号(如á、é、í、ó、ö、ő、ú、ü、ű)。文档类型包括发票(带有项目表格、增值税计算和银行详细信息)、正式信件(带有发件人/收件人标题和结构化正文)、表格(带有姓名、地址、税号、TAJ等标签字段)以及混合内容的自由格式文档(如标题、片段、列表、金额)。视觉上具有可变性:使用随机系统字体(TTF/OTF)和字体大小(10-22磅),DPI范围为150-300,可选高斯噪声(30%概率,严重程度0.02-0.08)和轻微旋转(±3°,20%概率),并支持两种渲染后端(Pillow默认,ReportLab+ImageMagick用于结构化发票)。文本内容包含真实的匈牙利名称、地址、城市、街道名称,有效格式的税号(HUxxxxxxxx)、IBAN/BIC代码、发票号码,以及福林货币格式和领域特定词汇(产品、服务、法律短语)。数据字段包括图像路径列表和文本字符串(以换行符分隔)。数据分割为单个分割,其中约60%为基于模板的文档(约1800个样本),约40%为自由格式文档(约1200个样本)。创建过程涉及从策划列表采样内容、使用预定义模板或随机组合组装布局,并通过PIL/Pillow渲染图像,可选后处理。使用注意事项:数据为合成生成,缺乏真实世界扫描伪影(如倾斜、阴影、纸张纹理),需考虑领域适应;专注于匈牙利语,保留变音符号和正字法;所有个人数据均为合成生成,不对应真实个体。限制包括字体覆盖依赖系统可用字体、布局多样性有限(模板样本结构固定,自由格式样本可能未覆盖所有真实布局),且所有文本为机器渲染,不适用于手写文本识别。
This dataset is a synthetically generated Hungarian document OCR dataset, containing 3000 Hungarian document images and their corresponding ground truth text transcriptions. It is specifically designed for training and evaluating optical character recognition (OCR), document layout analysis, and multimodal language models on Hungarian administrative and commercial documents. The dataset language is Hungarian, fully supporting diacritics (e.g., á, é, í, ó, ö, ő, ú, ü, ű). Document types include invoices (with item tables, VAT calculations, and bank details), formal letters (with sender/recipient headers and structured body), forms (with labeled fields such as name, address, tax number, TAJ), and free-form documents with mixed content (e.g., headings, paragraphs, lists, amounts). Visually, it exhibits variability: random system fonts (TTF/OTF) and font sizes (10-22 points) are used, DPI ranges from 150-300, optional Gaussian noise (30% probability, severity 0.02-0.08) and slight rotation (±3°, 20% probability) are applied, and two rendering backends are supported (Pillow by default, ReportLab+ImageMagick for structured invoices). Text content includes realistic Hungarian names, addresses, cities, street names, validly formatted tax numbers (HUxxxxxxxx), IBAN/BIC codes, invoice numbers, as well as Hungarian forint currency formats and domain-specific vocabulary (products, services, legal phrases). Data fields consist of a list of image paths and text strings (separated by newlines). The data is split into a single split, with approximately 60% being template-based documents (about 1800 samples) and approximately 40% free-form documents (about 1200 samples). The creation process involves sampling content from curated lists, assembling layouts using predefined templates or random combinations, and rendering images via PIL/Pillow, with optional post-processing. Usage notes: The data is synthetically generated and lacks real-world scanning artifacts (e.g., skew, shadows, paper texture), requiring consideration for domain adaptation; it focuses on Hungarian, preserving diacritics and orthography; all personal data is synthetically generated and does not correspond to real individuals. Limitations include font coverage dependent on system-available fonts, limited layout diversity (template samples have fixed structures, and free-form samples may not cover all real layouts), and all text is machine-rendered, making it unsuitable for handwritten text recognition.




