himalaya-ai/nepalipixel-synthetic-ocr-benchmark
收藏资源简介:
output_benchmark目录包含一个使用Nepali Pixel流水线生成的尼泊尔语(天城文)脚本的合成OCR基准数据集。该数据集设计用于评估OCR模型,具有最小的增强噪声和完全渲染的页面。数据集包括约15,000个图像-文本对,覆盖五个粒度级别:单词、句子、段落、页面和精确级别。图像格式为PNG灰度图,每个样本在metadata.jsonl中都有JSON条目,包含id、image_path、text、level、font_name、size_px、augmentations、intensity、source、image_w和image_h等字段。数据源为himalaya-ai/cc100-nepali,生成时使用benchmark模式以避免激进增强和裁剪,确保干净的基准图像。
The output_benchmark directory houses a synthetic OCR benchmark dataset of Nepali (Devanagari script) generated via the Nepali Pixel pipeline. This dataset is engineered for OCR model evaluation, featuring minimal augmented noise and fully rendered pages. It contains roughly 15,000 image-text pairs spanning five granularity tiers: word, sentence, paragraph, page, and exact level. All images are in PNG grayscale format, and each sample has a corresponding JSON entry in metadata.jsonl, including fields such as id, image_path, text, level, font_name, size_px, augmentations, intensity, source, image_w, and image_h. The data source is himalaya-ai/cc100-nepali, and the dataset was generated in benchmark mode to avoid aggressive augmentation and cropping, ensuring clean benchmark images.




