EPHOIE
收藏资源简介:
EPHOIE是首个针对中文的OCR和视觉信息提取的公开数据集,由华南理工大学创建。该数据集包含1494张来自中国各学校真实考试试卷头部的扫描图像,总计15,771个手写或打印的中文文本实例。数据集的创建过程涉及从考试试卷中裁剪出包含所有关键信息的头部区域。EPHOIE的应用领域主要集中在文档智能和视觉信息提取,旨在解决复杂布局和背景下的文本检测、识别及信息提取问题。
EPHOIE is the first publicly available dataset for Chinese-oriented optical character recognition (OCR) and visual information extraction, developed by South China University of Technology. This dataset contains 1,494 scanned images of the header sections of real exam papers from schools across China, with a total of 15,771 handwritten or printed Chinese text instances. The dataset construction process involves cropping the header regions containing all critical information from the exam papers. The application scenarios of EPHOIE mainly focus on document intelligence and visual information extraction, aiming to address the challenges of text detection, recognition and information extraction under complex layouts and backgrounds.

- 1Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution华南理工大学 · 2021年



