遇见数据集

HalalBench: A Multilingual OCR Benchmark Dataset for Food Packaging Ingredient Extraction

收藏
Zenodo2026-02-17 更新2026-05-26 收录
官方服务:

资源简介:

HalalBench is the first large-scale, multilingual benchmark for evaluating OCR engines on real-world food packaging images. The dataset comprises 1,043 images (50 real, 993 synthetic) with 36,438 annotations in COCO format spanning 14 languages: Arabic, Danish, Dutch, English, French, German, Indonesian, Japanese, Korean, Malay, Norwegian, Swedish, Thai, and Turkish. No existing OCR benchmark targets the unique challenges of ingredient labels: curved surfaces, dense multilingual text, and sub-8pt fonts. HalalBench fills this gap. Key results: No OCR engine exceeds 0.55 fuzzy F1, confirming food packaging OCR remains unsolved. ML Kit leads at 0.487 exact F1, followed by docTR (0.465), EasyOCR (0.210), and RapidOCR (0.189). Created by HalalLens (https://halallens.no), a production halal food app and halal scanner serving 20+ countries. Dataset and benchmark code: https://github.com/halallens-no/halalbench

提供机构:
Zenodo
创建时间:
2026-02-17
二维码
社区交流群
二维码
科研交流群
商业服务