遇见数据集

FieldBench Corpus: A Cross-Domain Benchmark for Schema-Driven Document Extraction

收藏
Zenodo2026-07-29 更新2026-08-01 收录
官方服务:

资源简介:

A cross-domain, field-level benchmark for schema-driven document extraction (document → structured JSON): 1,442 documents across 10 categories with per-field ground truth, released so extraction-accuracy claims become falsifiable and comparable. 769 real documents (SEC EDGAR, Caselaw Access Project, MTSamples, SROIE, U.S. government forms) and 673 synthetic documents, each labeled by source. Five categories are dual-source (matched-pair): real documents plus synthetic ones generated against the same schema, so the synthetic-vs-real accuracy gap can be measured holding category constant. A datasheet documents composition, provenance, and per-source licensing. Read before quoting any accuracy number: ~90% of documents are extraction-from-clean-text rather than rendered-page extraction, and synthetic documents overestimate accuracy relative to real ones — always report results stratified by source. See the datasheet in the repository. The scorer and run harness are at https://github.com/fieldbench/fieldbench.

提供机构:
Zenodo
创建时间:
2026-07-24
二维码
社区交流群
二维码
科研交流群
商业服务