arcolab-dev/FinDoc-Robust
收藏资源简介:
FinDoc-Robust是一个多模态、基准级的数据集,专为文档布局分析、视觉信息提取以及评估模型在真实世界退化情况下的鲁棒性而设计。该数据集包含5类不同的财务报告文档(例如现金流量表、资产负债表、试算平衡表、股东权益表、公司利润表)。对于每个文档,它提供了完美的数字向量、表格真值、像素级边界框,以及5个结构退化的变体,模拟相机捕捉、扫描和物理伪影。关键应用包括:鲁棒文档AI(训练模型抵抗几何失真、噪声和模糊)、表格重建(基准测试端到端的图像到Excel/HTML/Markdown管道)和多模态对齐(在复杂财务结构上微调如LayoutLMv3、Donut或专有视觉-LLMs等模型)。数据集结构按文档类型和数字索引分层组织,每个样本文件夹包含完整的模态子集,如PDF原始文件、XLSX目标布局、边界框JSON文件(包括向量和像素坐标)以及5个脏变体图像和对应边界框。
FinDoc-Robust is a multimodal, benchmark-grade dataset designed for Document Layout Analysis (DLA), Visual Information Extraction (VIE), and evaluating model robustness against real-world degradation. The dataset contains financial reports across 5 distinct document categories (e.g., cash flow statements, balance sheets, trial balances, shareholders equity, corporate income statements). For every document, it provides perfect digital vectors, tabular ground truths, pixel-level bounding boxes, and 5 structurally degraded (dirty) variants simulating camera captures, scans, and physical artifacts. Key applications include: Robust Document AI (training models to resist geometric distortions, noise, and blur), Table Reconstruction (benchmarking end-to-end Image-to-Excel/HTML/Markdown pipelines), and Multimodal Alignment (fine-tuning models like LayoutLMv3, Donut, or proprietary Vision-LLMs on complex financial structures). The dataset is organized hierarchically by document type and numerical index, with each sample folder containing a complete sub-set of modalities such as original PDF files, XLSX target layouts, bounding box JSON files (including vector and pixel coordinates), and 5 dirty variant images with corresponding bounding boxes.



