MultiOCR-QA
收藏资源简介:
MultiOCR-QA是一个多语言问答数据集,由因斯布鲁克大学创建,包含英语、法语和德语三种语言,共有60K个问题-答案对。数据集由经过OCR处理的历史文本构成,旨在评估OCR噪声对问答系统性能的影响。数据集涵盖了从OCR文本中直接生成的原始文本(RawOCR)和经过校正的文本(CorrectedOCR),使得可以在不同文本质量条件下直接比较问答性能。
MultiOCR-QA is a multilingual question answering dataset developed by the University of Innsbruck. It covers three languages: English, French and German, with a total of 60K question-answer pairs. The dataset is built from OCR-processed historical texts, and its core goal is to evaluate the impact of OCR-induced noise on the performance of question answering systems. It includes two types of texts: raw OCR-processed texts (RawOCR) and corrected texts (CorrectedOCR), which enables direct comparison of question answering performance under different text quality conditions.

- 1MultiOCR-QA: Dataset for Evaluating Robustness of LLMs in Question Answering on Multilingual OCR Texts因斯布鲁克大学 · 2025年



