BoundingDocs
收藏资源简介:
BoundingDocs是一个用于文档问答(Document Question Answering, QA)的统一数据集,包含空间注释。它整合了多个来自文档AI和视觉丰富文档理解(VRDU)领域的公共数据集,并将信息提取(IE)任务重新表述为QA任务。每个问答对通过边界框与其在文档中的位置相关联,增强了布局理解并减少了模型输出中的幻觉风险。数据集支持多种语言,包括英语、意大利语、西班牙语、法语、德语、葡萄牙语、中文和日语。数据集的结构包括文档来源、文档ID、文档图像、OCR结果以及问答对。数据集分为训练集、验证集和测试集,并提供了详细的统计信息。
BoundingDocs is a unified dataset for Document Question Answering (QA) with spatial annotations. It integrates multiple public datasets from the fields of Document AI and Visual Rich Document Understanding (VRDU), and reformulates Information Extraction (IE) tasks as QA tasks. Each QA pair is associated with its position in the document via a bounding box, which enhances layout understanding and reduces the risk of hallucinations in model outputs. The dataset supports multiple languages including English, Italian, Spanish, French, German, Portuguese, Chinese and Japanese. The dataset structure includes document source, document ID, document image, OCR results and QA pairs. The dataset is split into training, validation and test sets, with detailed statistical information provided.




