ServiceNow-AI/SynthDocBench
收藏资源简介:
SynthDocBench 是一个完全合成的基准数据集,用于评估视觉语言模型(VLMs)在复杂多页PDF文档上的性能。该数据集通过LLM管道端到端生成,包含具有嵌入D3.js图表、丰富布局和确定性基础真实答案的现实多页报告,支持无噪声、受控的评估。数据集包含三个子集:图表阅读(chart)、复杂多跳问答(complex)和跨模态问答(cross_modal),每个子集有不同难度级别(L1到L5)。数据集涵盖200个文档,总计513个问题,平均每个文档51页、17个图表和20,568个单词,包括20种图表类型和6种布局原型。数据模式包括问题、答案、难度、问题类型等字段,并基于JSON-LD元数据生成,确保答案的确定性和可追溯性。主题覆盖AI与技术、科学、经济与社会、环境和医学与健康等领域。
SynthDocBench is a fully synthetic benchmark for evaluating vision-language models (VLMs) on complex, multi-page PDF documents. Documents are generated end-to-end by an LLM pipeline that produces realistic multi-page reports with embedded D3.js charts, rich layouts, and deterministically grounded ground-truth answers — enabling controlled, noise-free evaluation. The dataset includes three subsets: chart reading, complex multi-hop QA, and cross-modal QA, each with difficulty levels from L1 to L5. It comprises 200 documents with 513 total questions, averaging 51 pages, 17 charts, and 20,568 words per document, covering 20 chart types and 6 layout archetypes. The data schema includes fields such as question, answer, difficulty, and question type, with answers derived from JSON-LD metadata for deterministic grounding. Topics span AI & Technology, Science, Economics & Society, Environment, and Medicine & Health.



