CC-OCR
收藏资源简介:
CC-OCR是由阿里巴巴集团和华中科技大学共同创建的综合性OCR基准数据集,旨在评估大型多模态模型在识字能力方面的表现。该数据集包含四个主要任务:多场景文本阅读、多语言文本阅读、文档解析和关键信息提取,涵盖39个子任务,包含7,058张全标注图像,其中41%来自实际应用。数据集的创建过程注重多样性、实用性和挑战性,涵盖自然场景、真实文档和手写图像等多种数据源。CC-OCR的应用领域广泛,包括文档数字化、办公机器人和城市监控等,旨在解决复杂文本识别和多模态理解的问题。
CC-OCR is a comprehensive OCR benchmark dataset jointly developed by Alibaba Group and Huazhong University of Science and Technology. It is designed to evaluate the text literacy performance of large multimodal models. This dataset includes four core tasks: multi-scene text reading, multi-language text reading, document parsing, and key information extraction, covering 39 subtasks, and contains 7,058 fully annotated images, 41% of which are sourced from real-world applications. The construction of CC-OCR prioritizes diversity, practicality and challenge, incorporating multiple data sources such as natural scene texts, real-world documents and handwritten images. CC-OCR has a wide range of application scenarios, including document digitization, office robotics, urban surveillance and so on, aiming to address the challenges of complex text recognition and multimodal understanding.

- 1CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy阿里巴巴集团 · 2024年



