chujian-ocr-shiwen-20260901
收藏资源简介:
本数据集为 OCR 释文三步流水线的交付结果,包含从古籍或文档图像中提取的识别文本。数据经过三步处理:首先使用百度 PaddleOCR-VL-1.5 进行保守首识,然后由 Qwen3-VL-4B-Instruct 对照页图和百度结果进行复核,最后使用 GPT-5.6 仅处理分歧、低置信和版式异常页,不对已一致页做语言润色。字段包括:`paddle_raw`(百度首识结果)、`qwen_review`(Qwen 复核结果)、`gpt_review`(GPT 处理结果)、`accepted_text`(最终接受文本)、`review_status`(审核状态)和 `source_sha256`(源文件 SHA-256 哈希)。对于古字图版、无码字和无法解释的分歧,必须放入 `quarantine` 目录,不得静默修改。资产详情、页数、SHA-256 与模型 revision 信息见 `HF_ASSET_MANIFEST_20260901.json` 文件。
This dataset is the delivery result of a three-step OCR transcription pipeline, containing recognized text extracted from ancient books or document images. The data undergoes three steps: first, conservative initial recognition using Baidu PaddleOCR-VL-1.5; second, review by Qwen3-VL-4B-Instruct against page images and Baidu results; third, GPT-5.6 is used only to handle discrepancies, low-confidence, and layout anomalies, without language polishing for consistent pages. Fields include: `paddle_raw` (Baidu initial recognition result), `qwen_review` (Qwen review result), `gpt_review` (GPT processing result), `accepted_text` (final accepted text), `review_status` (review status), and `source_sha256` (SHA-256 hash of source file). For ancient character plates, uncoded characters, and inexplicable discrepancies, they must be placed in the `quarantine` directory without silent modification. Asset details, page count, SHA-256, and model revision information are in the `HF_ASSET_MANIFEST_20260901.json` file.
数据集概述
该数据集为“chujian-ocr-shiwen-20260901”,是一个针对2026年9月1日正式释文OCR任务的本地交付根目录,围绕文档OCR识别与人工复核构建了一套三步流水线处理流程。
处理流水线
- 第一步(百度 PaddleOCR-VL-1.5):进行保守的初步识别;
- 第二步(Qwen3-VL-4B-Instruct):对照页面图像与百度识别结果进行复核;
- 第三步(GPT-5.6):仅处理存在分歧、低置信度或版式异常的页面,不对已达成一致的页面进行语言润色。
字段结构
每条识别结果按字段分离保存,具体包括:
paddle_raw:百度初识原始结果;qwen_review:Qwen复核结果;gpt_review:GPT处理结果;accepted_text:最终接受的文本;review_status:复核状态;source_sha256:源文件哈希值。
特殊处理规则
对于古字图版、无码字以及无法解释的分歧,必须归入 quarantine(隔离区),严禁静默修改为通顺文字。
附加文件说明
数据集中包含 HF_ASSET_MANIFEST_20260901.json 清单文件,其中记录了资产详情、页数、SHA-256哈希值以及所使用模型的修订版本号等信息。





