遇见数据集

LocalDoc/azerbaijani-htr-synthetic

收藏
Hugging Face2026-04-25 更新2026-05-03 收录
官方服务:

资源简介:

这是一个用于训练阿塞拜疆语手写文本识别模型的大规模合成数据集,专门针对阿塞拜疆语拉丁字母。数据集通过程序化流程生成,结合了真实手写字体和扫描风格增强,以模拟真实手写场景。它解决了阿塞拜疆语手写OCR数据公开可用性不足的问题,因为阿塞拜疆语是一种低资源语言,没有类似IAM的语料库。数据集包含约1,500,000个行图像,图像格式为JPEG,平均尺寸约800×70像素,总大小约30GB,分为训练集(95%)、验证集(2.5%)和测试集(2.5%)。每个样本包括图像(行级RGB图像)、文本(地面真实转录,NFC规范化)、字体(用于渲染的字体文件名)和配置文件(应用的增强配置文件,如mixed、school、office、archival)。生成过程涉及文本语料库组装(来自平行语料库和专门生成的字符串)、字体收集和验证(确保覆盖阿塞拜疆语特定字符并排除装饰性字体)、图像生成(包括单词级渲染、基线波动、墨水效果和背景增强)以及配置文件选择。数据集主要用于预训练阶段,以提高OCR模型在真实手写文档上的性能,但需注意数据是合成的,可能需结合真实数据进行微调。

A large-scale synthetic dataset for training handwritten text recognition (HTR) models on Azerbaijani Latin script. Generated using a procedural pipeline that combines real-world handwriting fonts with realistic scan-style augmentations. This dataset addresses the lack of publicly available Azerbaijani handwriting OCR data — a low-resource language for which no IAM-equivalent corpus exists. It contains approximately 1,500,000 line images in JPEG format, with an average size of about 800×70 pixels, totaling around 30 GB. The dataset is split into train (95%), validation (2.5%), and test (2.5%) sets. Each sample includes an image (line-level RGB image of synthesized handwriting), text (ground truth transcription, NFC-normalized), font (font filename used for rendering), and profile (augmentation profile applied, such as mixed, school, office, archival). The generation process involves text corpus assembly (from parallel corpus and specialized strings), font collection and validation (ensuring coverage of Azerbaijani-specific characters and excluding decorative fonts), image generation (including word-level rendering, baseline waviness, ink effects, and backgrounds), and profile selection. The dataset is intended for pretraining stages to improve OCR model performance on real handwritten documents, with caveats about synthetic nature and potential need for fine-tuning on real data.

提供机构:
LocalDoc
二维码
社区交流群
二维码
科研交流群
商业服务