Clouda OCR Arabic OCR Benchmark v0.1.0 — 177-Page Reproducible Evaluation
收藏资源简介:
Clouda OCR Arabic OCR Benchmark v0.1.0 is a reproducible evaluation of OCR and vision-language models on 177 distorted Arabic document pages. The benchmark was completed and first published on September 7, 2026. It evaluates Arabic text recognition quality using CER, WER, and normalized CER, with stored outputs and reproducibility artifacts. Models evaluated include HunyuanOCR-1.5, MBZUAI/AIN-7B, Qari OCR 0.4.0, Qwen3-VL-4B-Instruct, DeepSeek-OCR-2, and Arabic Nougat Large. Additional model attempts and partial results are documented separately. On this benchmark set, HunyuanOCR-1.5 achieved the lowest normalized Arabic CER. This record is intended to provide a persistent, citable reference for the benchmark methodology, evaluation results, and reproducibility materials associated with Clouda OCR. Project website:https://cloudaocr.xyz Source repository:https://github.com/sahrasayed3-crypto/clouda-ocr-benchmark



