pleias-post-ocr-correction-chonkie-aligned-en
收藏资源简介:
PleIAs后OCR校正—Chonkie对齐语义分块数据集是一个语义分块且跨度对齐的衍生数据集,源自PleIAs/Post-OCR-Correction原始数据集。该数据集专为OCR后校正、噪声文本规范化、历史文档处理及相关系统评估任务而设计。数据内容包含从原始text字段提取的OCR假设文本块,以及从corrected_text字段提取的对应后OCR校正输出块,同时继承了PleIAs数据集的元数据。每个记录还包含将分块链接回原始源文档的字符跨度信息,以及在过滤过程中生成的对齐诊断数据。需要特别注意的是,原始PleIAs数据集中的corrected_text字段是实验性的多语言后OCR校正输出,不应被视为经过人工验证的真实数据。因此,在本数据集中,字段映射为:ocr_hypothesis对应原始PleIAs的text,ground_truth对应PleIAs的corrected_text。数据集经过了严格的过滤处理,使用自动启发式方法移除了可疑的对齐案例,包括空OCR或校正块、非常短的分块、极端的OCR/校正长度比例、低字符级相似度以及高CER类编辑距离等情况。从检查的86,469条记录中,保留了52,962条,移除了33,507条(占比38.75%)。数据以JSONL格式存储,每条记录包含文档元数据、真实文本和OCR假设文本等字段。由于校正目标是合成的,该数据集更适合用于弱监督学习、预训练、过滤实验或诊断分析,而非作为黄金标准的基准评估。数据集支持英语、法语、意大利语和德语,主要涉及文本生成任务。
PleIAs Post-OCR Correction – Chonkie-Aligned Semantic Chunking Dataset is a semantic chunking and span-aligned derived dataset originating from the original PleIAs/Post-OCR-Correction dataset. This dataset is specifically designed for tasks including OCR post-correction, noisy text normalization, historical document processing, and related system evaluation. The dataset content includes OCR hypothesis text chunks extracted from the original "text" field, and corresponding post-OCR correction output chunks extracted from the "corrected_text" field, while inheriting the metadata from the PleIAs dataset. Each record also contains character span information that links chunks back to the original source document, as well as alignment diagnostic data generated during the filtering process. It is particularly important to note that the "corrected_text" field in the original PleIAs dataset is an experimental multilingual post-OCR correction output and should not be treated as manually verified ground truth data. Therefore, the field mapping in this dataset is as follows: "ocr_hypothesis" corresponds to the "text" field from the original PleIAs dataset, and "ground_truth" corresponds to the "corrected_text" field from PleIAs. The dataset underwent rigorous filtering processing, where automatic heuristic methods were used to remove suspicious alignment cases, including empty OCR or correction chunks, extremely short chunks, extreme OCR/correction length ratios, low character-level similarity, and high CER (Character Error Rate)-based edit distance, among other scenarios. Out of the 86,469 inspected records, 52,962 were retained and 33,507 were removed, accounting for 38.75% of the total. The data is stored in JSONL format, with each record containing fields such as document metadata, ground truth text, and OCR hypothesis text. Since the correction targets are synthetic, this dataset is more suitable for weakly supervised learning, pre-training, filtering experiments, or diagnostic analysis, rather than serving as a gold-standard benchmark evaluation. The dataset supports English, French, Italian, and German, and primarily relates to text generation tasks.




