NuosuBburma-OCR-Evaluation-Set
收藏资源简介:
NuosuBburma OCR Evaluation Set 是一个专门用于评估规范彝文(Nuosu Yi)光学字符识别(OCR)模型性能的基准数据集,尤其适配PaddleOCR-VL风格的图片OCR评测任务。数据集旨在全面测试模型在多样化、真实世界彝文文本场景下的识别能力。数据内容与构成方面,数据集共包含758个样本,每个样本由一张图片和对应的唯一标准真值(ground truth)文本组成。图片存储于`images/`目录,真值文本及元数据记录在`annotations.jsonl`文件中,每行对应一个样本,采用结构化的JSON格式。该格式不仅包含图片路径和真值文本,还通过`meta`字段提供了丰富的样本属性信息,包括来源名称、样本类型(单行、区域或整页)、采集场景(如新书PDF、旧书PDF、真实屏幕照片、工整手写照片、真实场景照片)、难度等级(简单、中等、困难)以及文字混合类型(纯彝文、彝汉混排、彝汉拉丁注音、纯汉文等)。数据规模为758个图文对,确保了数据集的完整性,标注中不存在空真值文本或缺失图片的情况。数据来源和类型高度多样化,覆盖了旧书扫描、新书扫描、屏幕截图拍照、手写文本拍照、真实环境照片以及标牌照片等多种场景。文本内容不仅包括纯彝文,还涵盖了彝汉混排文本、带有拉丁字母注音的彝文、数字、页码、脚注以及封面和工具书条目,从而全面评估OCR系统处理复杂、混合文本的能力。该数据集适用于OCR模型的性能评估与基准测试任务,特别是针对少数民族语言(彝文)以及多语言混合场景下的文本识别研究。
The NuosuBburma OCR Evaluation Set is a benchmark dataset specifically designed for evaluating the performance of optical character recognition (OCR) models for standardized Yi script (Nuosu Yi), particularly adapted for PaddleOCR-VL style image OCR evaluation tasks. The dataset aims to comprehensively test the models recognition capabilities in diverse, real-world Yi text scenarios. In terms of data content and structure, the dataset contains a total of 758 samples, each consisting of an image and its corresponding unique ground truth text. Images are stored in the `images/` directory, while ground truth text and metadata are recorded in the `annotations.jsonl` file, with each line corresponding to a sample in a structured JSON format. This format includes not only image paths and ground truth text but also provides rich sample attribute information through the `meta` field, such as source name, sample type (single-line, region, or full-page), collection scenario (e.g., new book PDF, old book PDF, real screen photos, neat handwriting photos, real scene photos), difficulty level (easy, medium, hard), and text mixing types (pure Yi, Yi-Chinese mixed, Yi-Chinese with Latin phonetic notation, pure Chinese, etc.). The data scale consists of 758 image-text pairs, ensuring the completeness of the dataset with no empty ground truth texts or missing images in the annotations. Data sources and types are highly diverse, covering various scenarios such as old book scans, new book scans, screen screenshot photos, handwritten text photos, real environment photos, and signboard photos. The text content includes not only pure Yi script but also Yi-Chinese mixed texts, Yi with Latin alphabet phonetic notation, numbers, page numbers, footnotes, as well as covers and reference book entries, thereby comprehensively evaluating the OCR systems ability to handle complex and mixed text. The dataset is suitable for performance evaluation and benchmark testing of OCR models, especially for research on minority languages (Yi script) and multilingual mixed text scenarios.
数据集概述
数据集名称:NuosuBburma OCR Evaluation Set(规范彝文OCR评估集)
语言:彝文(ii)、中文(zh)
任务类型:图像到文本(image-to-text)
数据集规模:样本数小于1000(n<1K)
许可证:其他(other)
标签:OCR、PaddleOCR、Nuosu Yi、Yi、evaluation
核心内容
-
核心文件:
annotations.jsonl:唯一的标准真实数据文件,每行包含图片路径和真实文本。images/:评估图片文件夹,路径与annotations.jsonl中的images字段一致。
-
数据规模:
- 样本数:758
- 图片数:758
- 空真实文本:0
- 缺图:0
覆盖类型
| 维度 | 覆盖范围 |
|---|---|
| 图片来源 | 旧书扫描、新书扫描、屏幕拍照、手写拍照、真实场景照片、标牌照片 |
| 输入粒度 | 单行(line)、区域(region)、整页(page) |
| 文字类型 | 纯彝文、彝汉混排、彝汉拉丁注音、数字、页码、脚注、封面和工具书条目 |
| 难度 | easy / medium / hard |
标注格式
annotations.jsonl 每行包含以下字段:
- id:样本ID
- images:图片路径列表
- messages:对话格式,包含用户输入(如
<image>OCR:)和助手输出(真实文本) - meta:元数据,包括:
source_name:来源名称sample_type:样本类型(line / region / page)scene:场景类型(new_print_pdf / old_print_pdf / real_screen_photo / neat_handwriting / real_photo)difficulty:难度(easy / medium / hard)script_mix:文字混合情况(yi / yi_han / yi_han_latin / han / other)
版本信息
- 最终版本:
20260629_final_locked_gt_complete - 最终锁定时间:
2026-06-29T21:22:14 annotations.jsonlSHA256:db9a2a061054ccfa772566c4482aa0793a8df72b49888d251e0de0e5eb78be93





