ByteDance/AncientDoc
收藏资源简介:
AncientDoc是一个专为中文古籍文档理解设计的全面基准数据集。它包含从OCR到知识推理的多任务评估,旨在推动多模态大型模型在古籍场景下的识别、理解和推理能力的研究。数据集包含2973页文献,涵盖从战国时期到清朝的多个重要历史时期,分为14类文献类型。任务类型包括页面级OCR、文言文翻译、基于推理的问答、基于知识的问答和语言变体问答。
AncientDoc is the first comprehensive benchmark dataset specifically designed for Chinese Ancient Document Understanding. It covers multi-task evaluation ranging from OCR to knowledge reasoning, aiming to promote research on the recognition, understanding, and reasoning capabilities of multimodal large models in the scenario of ancient documents. The dataset contains 2,973 pages of literature, spanning from the Warring States period to the Qing Dynasty, divided into 14 categories of literature types. Task types include page-level OCR, vernacular translation, reasoning-based QA, knowledge-based QA, and linguistic variant QA.




