himalaya-ai/nepali-deva-ocr-eval
收藏资源简介:
该数据集名为“Nepali Deva Eval”,是一个用于评估尼泊尔语Devanagari OCR(光学字符识别)质量的保留切片数据集,专门针对直接尼泊尔语提取任务。它属于图像到文本任务类别,支持尼泊尔语、印地语和马拉地语,并包含OCR、Devanagari和glm-ocr相关标签。数据集的核心内容包括id(唯一样本标识符)、image(图像文件的相对路径)和ocr(真实文本标签)。数据来源100%来自“nepali_ocr_printed”默认训练集,主要用于微调和评估,例如通过GLM OCR工作空间生成。主要原始文件为*.ocr.jsonl格式,包含图像、OCR结果、来源仓库和语言/来源列;可选文件为*.sharegpt.json格式,包含消息和图像。该数据集旨在帮助评估OCR模型在尼泊尔语Devanagari文本提取中的性能。
The dataset is named Nepali Deva Eval and serves as a held-out slice for evaluating the quality of Nepali Devanagari OCR (Optical Character Recognition) in direct Nepali extraction tasks. It falls under the image-to-text task category, supports languages including Nepali, Hindi, and Marathi, and is tagged with OCR, Devanagari, and glm-ocr. Core columns include id (unique sample identifier), image (relative path to the image file), and ocr (ground-truth text label). The data is sourced 100% from the nepali_ocr_printed default training split. It is generated from the GLM fine-tuning workspace and is intended for fine-tuning and evaluation purposes, with main raw files in *.ocr.jsonl format (containing image, ocr, source_repo, and language/provenance columns) and optional files in *.sharegpt.json format (with messages and images). The dataset aims to assess OCR model performance in extracting Nepali Devanagari text.




