OCR-Tibetan_line_to_text_benchmark
收藏资源简介:
该数据集是一个用于评估和比较藏文OCR模型的基准数据集。它包含了多种脚本、书写风格和印刷方法,能够全面测试OCR模型在不同领域的表现。数据集的特征包括文件名、标签、图像URL、BDRC工作ID、字符长度、脚本类型、书写风格和印刷方法。数据集被分为多个部分,每个部分代表不同的来源和风格,如Norbuketaka、Lithang_Kanjur等。数据集的总大小约为161MB,包含约496,000个示例。
This dataset is a benchmark dataset for evaluating and comparing Tibetan OCR models. It encompasses diverse scripts, writing styles and printing methods, enabling comprehensive testing of OCR models' performance across various domains. The features of the dataset include file names, labels, image URLs, BDRC work IDs, character lengths, script types, writing styles and printing methods. The dataset is divided into multiple sections, each representing distinct sources and styles such as Norbuketaka, Lithang_Kanjur, etc. The total size of the dataset is approximately 161 MB, containing around 496,000 examples.




