pleias-post-ocr-correction-chonkie-aligned-fr
收藏资源简介:
PleIAs后OCR校正 — Chonkie对齐语义分块数据集是从PleIAs/Post-OCR-Correction数据集衍生而来的语义分块和跨对齐版本。该数据集旨在支持OCR后校正、噪声文本规范化和历史文档处理等任务。每个数据记录包含一个来自原始text字段的OCR假设块、一个来自corrected_text字段的对应后OCR校正输出块、从PleIAs数据集继承的元数据、将每个块链接回原始源文档的字符跨度,以及在过滤过程中产生的对齐诊断信息。重要说明:原始PleIAs数据集中的corrected_text字段是实验性的多语言后OCR校正输出,不应被视为手动验证的真实数据;因此,在该衍生数据集中,ocr_hypothesis存储原始PleIAs文本,ground_truth存储PleIAs的corrected_text。数据集经过过滤,移除了可疑对齐案例(包括空块、非常短的块、极端OCR/校正长度比、低字符级相似性和高CER类编辑距离),检查了7,435条记录,保留了4,835条,移除了2,600条可疑记录(移除比例34.97%)。数据格式为JSONL,每个JSON记录包含document_metadata、ground_truth、ocr_hypothesis等字段。该数据集适用于OCR后校正、噪声文本规范化、历史文档处理以及后OCR校正系统评估的实验,但由于校正目标是合成的,更适合弱监督、预训练、过滤实验或诊断分析,而不是黄金标准基准评估。数据集支持英语、法语、意大利语和德语,许可证未知,任务类别为文本生成,标签包括OCR、后OCR校正、历史文档、语义分块、chonkie和合成。
PleIAs Post-OCR Correction — Chonkie Aligned Semantic Chunking Dataset is a semantically chunked and cross-aligned derivative of the PleIAs/Post-OCR-Correction dataset. This dataset is designed to support tasks including OCR post-correction, noisy text normalization, and historical document processing. Each data record contains an OCR hypothesis chunk from the original `text` field, a corresponding post-OCR correction output chunk from the `corrected_text` field, metadata inherited from the PleIAs dataset, character spans linking each chunk back to its original source document, and alignment diagnostic information generated during the filtering process. Important note: The `corrected_text` field in the original PleIAs dataset is an experimental multilingual post-OCR correction output and should not be treated as manually validated ground truth. Accordingly, in this derivative dataset, `ocr_hypothesis` stores the raw PleIAs text, while `ground_truth` stores the `corrected_text` from PleIAs. The dataset was filtered to remove suspicious alignment cases, including empty chunks, extremely short chunks, extreme OCR/correction length ratios, low character-level similarity, and high CER-style edit distance. A total of 7,435 records were examined, with 4,835 retained and 2,600 suspicious records removed (removal rate: 34.97%). The dataset is stored in JSONL format, with each JSON record containing fields such as `document_metadata`, `ground_truth`, and `ocr_hypothesis`. This dataset is suitable for experiments on OCR post-correction, noisy text normalization, historical document processing, and post-OCR correction system evaluation. However, since the correction targets are synthetic, it is more suitable for weakly supervised learning, pre-training, filtering experiments, or diagnostic analysis, rather than gold-standard benchmark evaluation. The dataset supports English, French, Italian, and German languages. Its license is unknown, the task category is text generation, and the tags include OCR, post-OCR correction, historical documents, semantic chunking, chonkie, and synthetic.




