遇见数据集

CGRE Framework Dataset - A Dataset automatically generated to evaluate OCR Software on Webdocuments

收藏
Zenodo2020-08-06 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Description</strong><br> The provided dataset was generated by the CGRE Framework.<br> It was generated as a part of a bachelor thesis and used to evaluate the Tesseract OCR Software on webdocuments. <strong>CGRE_dataset.zip:</strong><br> <em>1. crawl.json</em><br> This file contains crawling results from the alexa.com Top 50 most used webpages in the US from the 7th June 2020.<br> The crawling was done specifically for styling information only. <em>2. html</em><br> The generated webdocuments can be found in this directory.<br> They are based on the crawled styling information.<br> The levels of the directory are used to store the different styling attributes.<br> Every directory is named by the used value for a specific styling attribute.<br> Every word is placed in a span html element. <em>3. dataset</em><br> The rendered webdocuments can be found in this directory as png files.<br> They were rendered using the Chromium Embedded Framework (CEF) and contain corresponding labels.<br> The labels are in the same directory with the same name as the corresponding rendered webdocument, just as txt files.<br> The labels contain "word\t(left,top,width,height)\n" lines.<br> "(left,top,width,height)" is the bounding box of a span element containing a word.<br> "word" is the word in the bounding box. <em>4. dataset_tesseract_complete</em><br> This directory contains the Tesseract results on the dataset as txt files.<br> The structure is analogue to the dataset.<br> The txt files contain analogue to the dataset "word\t(left,top,width,height)\n" lines. <em>5. evaluation</em><br> The results of the evaluation of Tesseract on the dataset.<br> To evaluate the localisation of words by Tesseract, the Intersection Over Union metric was used, with different threshold values (0.5, 0.6, 0.7, 0.8, 0.9).<br> To evaluate the determination of words by Tesseract, a normalized Levenshtein distance metric was used, with different threshold values (0.5, 0.6, 0.7, 0.8, 0.9).<br> The times were measured by using this system:<br> Ubuntu 20.04, AMD Ryzen 5 1600 CPU, AMD Radeon RX Vega 56 GPU, 16 GB DDR4 RAM with 2400 MHz<br> The different threshold values are stored in the filenames.<br> You can find the results in the csv files.<br> Every line contains the results for a specific webdocument.<br> The txt files contain calculated precision and recall values.

提供机构:
Zenodo
创建时间:
2020-08-06
二维码
社区交流群
二维码
科研交流群
商业服务