OCR-IDL
收藏资源简介:
OCR-IDL数据集是由计算机视觉中心UAB创建,包含2600万页的工业文档,使用商业OCR引擎进行标注,具有超过20,000美元的估算价值。该数据集涵盖了化学、医学和药物行业的多种文档类型,如信件、报告、电子邮件等,时间跨度长达100年。创建过程中,通过Amazon Textract进行文档的预处理和标注,确保了高质量的文本和布局信息。OCR-IDL旨在推动文档智能领域的研究,特别是在自动文档结构化、分类和信息提取方面,以减少人工处理的时间和成本。
The OCR-IDL dataset was developed by the Computer Vision Center at UAB. It contains 26 million pages of industrial documents annotated using commercial OCR engines, with an estimated value exceeding $20,000. The dataset covers a wide range of document types from the chemical, medical and pharmaceutical industries, including letters, reports, emails and other types, with a time span of up to 100 years. During its creation, Amazon Textract was utilized for document preprocessing and annotation, ensuring high-quality text and layout information. The OCR-IDL dataset aims to advance research in the field of document intelligence, particularly in the areas of automatic document structuring, classification and information extraction, to reduce the time and cost associated with manual document processing.




