遇见数据集

huggingworld/ParseBench

收藏
Hugging Face2026-04-26 更新2026-05-03 收录
官方服务:

资源简介:

ParseBench是一个用于评估AI代理文档解析系统的基准测试数据集,专注于真实世界企业文档。数据集包含约2000页人工验证的标注页面,来自1200多份公开文档,涵盖保险、金融、政府等领域,文档难度从简单到对抗性困难不等。数据集分为五个能力维度:表格(评估合并单元格和层次化表头的结构保真度)、图表(从条形图、折线图、饼图等中精确提取数据点)、内容忠实度(检测遗漏、幻觉和阅读顺序违规)、语义格式化(保留具有语义的内联格式,如删除线、上标/下标、粗体等)和视觉定位(追踪每个提取元素在页面上的精确源位置)。总共包含超过169,000条测试规则,提供细粒度的诊断能力。数据格式为JSONL,每个测试规则一行,包含pdf路径、类别、ID、类型、规则、页码、预期markdown和标签等字段。数据集还附带完整的评估框架,支持端到端管道评估、每维度评分和跨管道比较。

ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation across five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics. The evaluation set contains ~2,000 human-verified pages from over 1,200 publicly available documents spanning insurance, finance, government, and other domains, ranging from straightforward to adversarially hard. It includes over 169K test rules across the five dimensions, providing fine-grained diagnostic power. All annotations are produced through a two-pass pipeline: frontier VLM auto-labeling followed by targeted human correction. The benchmark ships with a full evaluation framework supporting end-to-end pipeline evaluation, per-dimension scoring, and cross-pipeline comparison. The dataset is in JSONL format with one line per test rule, including fields such as pdf path, category, id, type, rule, page, expected_markdown, and tags.

提供机构:
huggingworld
二维码
社区交流群
二维码
科研交流群
商业服务