遇见数据集

thoughtworks/document-processing-benchmark

收藏
Hugging Face2026-05-24 更新2026-06-14 收录
官方服务:

资源简介:

Document Processing Benchmark是一个文档处理基准数据集,整合了8个公共文档数据集(包括收据、发票、表单、银行对账单、多页文档、合同等),并统一规范为Parquet格式。每个数据行包含文档、真实标注以及来自一个或多个参考模型的实时API调用的每行令牌数、延迟和成本数据。用户可以直接读取目标模型的成本、延迟和质量指标,无需重新运行模型。数据集提供多个配置(如test、full、v2、v3),针对不同生产工作负载场景(如单页文档全结构化提取、多页文档或窄字段提取、预处理路径比较等)。每行包含22个基础列(如文档ID、来源数据集、文档类型、图像base64编码、真实标注JSON等)和基于模型变体的基准列。数据集来源多样,许可证混合了宽松和研究专用类型,需根据每行的许可证列进行过滤以符合使用需求。

Document Processing Benchmark is a benchmark dataset for document processing that normalizes 8 public document datasets (receipts, invoices, forms, bank statements, multi-page docs, contracts) into one parquet schema. Each row has the document, ground-truth annotations, and per-row token/latency/cost numbers from real API calls to one or more reference models. You can read off a targets cost/latency/quality without re-running it. The dataset offers multiple configs (e.g., test, full, v2, v3) targeting different production workload shapes, such as single-page full structured extraction, multi-page or narrow header extraction, and preprocessing path comparisons. Every row includes 22 base columns (e.g., doc_id, source_dataset, doc_type, image_b64, ground_truth_json) and baseline columns following a naming convention for model variants. Sources are diverse with a mix of permissive and research-only licenses, requiring filtering by the per-row license column for appropriate use cases.

提供机构:
thoughtworks
二维码
社区交流群
二维码
科研交流群
商业服务