遇见数据集

PatentME Datasets for Printed Mathematical Expression Recognition

收藏
Zenodo2026-06-18 更新2026-06-18 收录
官方服务:

资源简介:

Overview The PatentME datasets accompany the paper: PatentME: A Dataset and Reference-Free Post-OCR Verification Task for Printed Mathematical Expression Recognition accepted for publication at ICDAR 2026. Preprint : https://hal.science/hal-05660080 PatentME provides large-scale collections of mathematical expressions extracted from European patent documents together with their associated MathML annotations. Unlike most existing Mathematical Expression Recognition (MER) datasets, which are generated from LaTeX source files of scientific publications, PatentME contains expressions originating from real patent documents exhibiting substantial variability in image quality, typography, scanning artifacts, layouts, and notation styles. The datasets are derived from publicly available data distributed by the European Patent Office (EPO). Why PatentME? Most publicly available MER datasets suffer from one or more of the following limitations: • Synthetic rendering from LaTeX sources.• Uniform fonts and layouts.• Minimal noise and degradation.• Expressions originating primarily from scientific publications.• Lack of realistic OCR challenges. PatentME addresses these limitations by providing: • Real-world mathematical expressions extracted from patent applications.• Large visual diversity in fonts, styles, resolutions, and document quality.• Ground-truth MathML annotations originating from industrial patent publication workflows.• Data suitable for OCR evaluation, training, robustness testing, and post-OCR verification. Dataset Collection The archive contains three complementary datasets. PatentME-OCR A benchmark dataset for mathematical expression recognition. Content • Approximately 40,000 mathematical expression images.• Corresponding MathML annotations from the EPO website.• Cleaned MathML annotations (see the MathML cleaning script on GitHub: https://github.com/fwieckowiak/PatentME).• Renderings of the cleaned MathML in both display and inline modes. Directory Structure PatentME-OCR/├── PatentME-OCR_data.csv├── PatentME-OCR_raw_img/├── PatentME-OCR_mml/├── PatentME-OCR_mml_cleaned/├── PatentME-OCR_mml_cleaned_display_img/└── PatentME-OCR_mml_cleaned_inline_img/ PatentME-OCR_raw_imgOriginal expression images extracted from patent documents PatentME-OCR_mmlOriginal MathML annotations PatentME-OCR_mml_cleanedNormalized MathML annotations PatentME-OCR_mml_cleaned_display_imgDisplay-mode rendering of cleaned MathML PatentME-OCR_mml_cleaned_inline_imgInline-mode rendering of cleaned MathML PatentME-Siamese A dataset designed for post-OCR verification. The objective is to determine whether a recognized mathematical expression exactly matches the original expression image without requiring access to the ground-truth annotation. To do so, this dataset contains positive and negative image pairs, where the positive pairs consist of the original patent expression image (which may contain noise, artifacts, and variability) and its corresponding rendering of the correct MathML, while the negative pairs consist of the original expression image and a rendering generated from an incorrect OCR prediction. Directory Structure PatentME-Siamese/├── PatentME-Siamese_pairs.csv├── PatentME-OCR_raw_img/├── PatentME-OCR_display_img/└── pred_texteller_display_img/ PatentME-600k A large-scale training dataset. Content • Approximately 600,000 mathematical expression images.• Corresponding MathML annotations. Directory Structure PatentME-600k/├── PatentME-600k_data.csv└── images/ MathML Cleaning Procedure The cleaned MathML annotations were generated through several normalization steps: • Whitespace normalization.• Unicode normalization.• Standardization of special characters.• Correction of formatting inconsistencies.• Replacement of deprecated MathML elements (e.g., ) with modern equivalents. The cleaned version is recommended for training and evaluation. Data Source The datasets originate from publicly available European Patent Office resources. Primary source: https://publication-bdds.apps.epo.org/raw-data/products/public/product/32 Earlier versions of PatentME-OCR were partially collected using: https://data.epo.org/expert-services/ The complete collection process is described in the accompanying publication and GitHub repository. Known Limitations • Some MathML annotations may contain residual formatting inconsistencies.• A small number of annotation errors may exist.• LaTeX annotations are not provided.• PatentME reflects the characteristics and biases of patent publications and may not generalize to all mathematical document types. Contributions, corrections, and bug reports are welcome. Acknowledgements PatentME is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. This product contains data sourced from EPO databases. © European Patent Organisation. Related projects and code : https://github.com/fwieckowiak/PatentME https://github.com/fwieckowiak/PiE-MER If these datasets contribute to a scientific publication, please cite the ICDAR 2026 paper (full citation pending). Please also cite this Zenodo record. This work was conducted as part of François Wieckowiak's PhD: https://liris.cnrs.fr/these/these-francois-wieckowiak carried out jointly with: • Luminess — https://www.luminess.eu/• LIRIS (Laboratoire d'Informatique en Image et Systèmes d'Information) — https://liris.cnrs.fr/ The authors thank the European Patent Office (EPO) for providing public access to patent publication data. For questions, bug reports, or dataset corrections, please contact François Wieckowiak: https://fwieckowiak.github.io/

提供机构:
Zenodo
创建时间:
2026-06-16
二维码
社区交流群
二维码
科研交流群
商业服务