CORU
收藏资源简介:
CORU数据集是由因斯布鲁克大学和DISCO公司联合开发的,旨在增强多语言环境下的OCR和收据信息提取能力。该数据集包含20,000条来自不同零售环境的收据,涵盖超市和服装店等多种场景。数据集创建过程中,通过严格的收集、标注和质量控制确保数据的多样性和真实性。CORU数据集特别适用于处理复杂和嘈杂的文档布局,如实际收据,并推动自动化多语言文档处理技术的发展。
The CORU dataset was co-developed by the University of Innsbruck and DISCO Corporation, aiming to enhance OCR and receipt information extraction capabilities in multilingual environments. This dataset contains 20,000 receipts from various retail settings, covering diverse scenarios such as supermarkets and clothing stores. Strict collection, annotation and quality control measures were adopted throughout the dataset development lifecycle to guarantee its diversity and authenticity. The CORU dataset is particularly suitable for handling complex and noisy document layouts such as real-world receipts, and promotes the advancement of automated multilingual document processing technologies.



