遇见数据集

Clouda OCR Arabic OCR Benchmark v0.1.0 — 177-Page Reproducible Evaluation

收藏
Zenodo2026-09-20 更新2026-10-01 收录
官方服务:

资源简介:

Clouda OCR Arabic OCR Benchmark v0.1.0 is a reproducible evaluation of OCR and vision-language models on 177 distorted Arabic document pages. The benchmark was completed and first published on September 7, 2026. It evaluates Arabic text recognition quality using CER, WER, and normalized CER, with stored outputs and reproducibility artifacts. Models evaluated include HunyuanOCR-1.5, MBZUAI/AIN-7B, Qari OCR 0.4.0, Qwen3-VL-4B-Instruct, DeepSeek-OCR-2, and Arabic Nougat Large. Additional model attempts and partial results are documented separately. On this benchmark set, HunyuanOCR-1.5 achieved the lowest normalized Arabic CER. This record is intended to provide a persistent, citable reference for the benchmark methodology, evaluation results, and reproducibility materials associated with Clouda OCR. Project website:https://cloudaocr.xyz Source repository:https://github.com/sahrasayed3-crypto/clouda-ocr-benchmark

提供机构:
Zenodo
创建时间:
2026-09-20
二维码
社区交流群
二维码
科研交流群
商业服务