CC-OCR
收藏资源简介:
CC-OCR是由阿里巴巴集团和华中科技大学共同创建的综合性OCR基准数据集,旨在评估大型多模态模型在识字能力方面的表现。该数据集包含四个主要任务:多场景文本阅读、多语言文本阅读、文档解析和关键信息提取,涵盖39个子任务,共有7058张全标注图像,其中41%来自实际应用场景。数据集的创建过程注重多样性、实用性和挑战性,涵盖自然场景、真实文档和手写图像等多种数据来源。CC-OCR的应用领域广泛,包括文档数字化、办公机器人和城市监控等,旨在解决复杂文本图像识别和理解的问题。
CC-OCR is a comprehensive OCR benchmark dataset jointly developed by Alibaba Group and Huazhong University of Science and Technology, aiming to evaluate the performance of large multimodal models in terms of literacy capabilities. This dataset encompasses four core tasks: multi-scene text reading, multi-language text reading, document parsing, and key information extraction, covering 39 subtasks and totaling 7058 fully annotated images, 41% of which are sourced from real-world application scenarios. The construction of CC-OCR emphasizes three core principles: diversity, practicality, and challenge, and incorporates diverse data sources including natural scenes, real documents, and handwritten images. CC-OCR has a wide range of application scenarios such as document digitization, office robotics, and urban surveillance, and is designed to solve the problems of complex text image recognition and understanding.




