遇见数据集

ENC-PSL/evahan-ultraglyph

收藏
Hugging Face2026-05-23 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是中文OCR行数据集,专门为LREC2026会议的EvaHan竞赛设计,用于中文文档的光学字符识别(OCR)和手写文本识别(HTR)任务,特别是历史文档处理。数据集包含行图像和对应的中文转录文本对,来源分为两类:第一类是fonts,使用中文字体合成的印刷体行,训练集有72,323条,验证集有12,676条;第二类是ultraglyph,使用真实数据和ultraglyph生成器合成的行,混合了手写和印刷体,训练集有57,063条,验证集有10,201条。总计训练集129,386条,验证集22,877条。转录文本长度范围从2到30个字符,中位数为16个字符。该数据集是ENCHANTeam团队为竞赛创建的合成数据集,团队在竞赛中最终排名第三。

This is a Chinese OCR line dataset specifically designed for the EvaHan competition at the LREC 2026 conference, targeting optical character recognition (OCR) and handwritten text recognition (HTR) tasks for Chinese documents, especially historical document processing. The dataset comprises pairs of line images and their corresponding Chinese transcriptions, which are divided into two categories: The first category is "fonts", which are printed line images synthesized using Chinese fonts. The training set contains 72,323 samples, and the validation set contains 12,676 samples. The second category is "ultraglyph", which are line images synthesized using real data and the ultraglyph generator, mixing handwritten and printed styles. The training set contains 57,063 samples, and the validation set contains 10,201 samples. In total, the training set has 129,386 samples and the validation set has 22,877 samples. The length of the transcribed text ranges from 2 to 30 characters, with a median of 16 characters. This dataset is a synthetic dataset created by the ENCHANTeam for the competition, and the team ultimately ranked third in the competition.

提供机构:
ENC-PSL
二维码
社区交流群
二维码
科研交流群
商业服务