KITAB-Bench
收藏资源简介:
KITAB-Bench是一个全面的阿拉伯OCR基准测试,由MBZUAI创建。该数据集包含8,809个样本,跨越9个主要领域和36个子领域,涵盖了多种文档类型,包括手写文本、结构化表格和针对商业智能的21种图表类型。数据集结合了来自现有数据集的精心挑选的样本、手动注释的PDFs以及通过LLM辅助管道生成的合成内容。该基准测试旨在为现代OCR系统提供严格的评估框架,并推动阿拉伯文档分析方法的改进。
KITAB-Bench is a comprehensive Arabic OCR benchmark created by MBZUAI. This dataset contains 8,809 samples spanning 9 major domains and 36 sub-domains, covering diverse document types including handwritten text, structured tables, and 21 types of charts for business intelligence. The dataset combines carefully selected samples from existing datasets, manually annotated PDFs, and synthetic content generated via LLM-augmented pipelines. This benchmark aims to provide a rigorous evaluation framework for modern OCR systems and drive advancements in Arabic document analysis methods.

- 1KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document UnderstandingMBZUAI · 2025年



