FieldBench Corpus: A Cross-Domain Benchmark for Schema-Driven Document Extraction
收藏资源简介:
A cross-domain, field-level benchmark for schema-driven document extraction (document → structured JSON): 1,114 documents across 12 categories with per-field ground truth, released so extraction-accuracy claims become falsifiable and comparable. 613 real documents (SEC EDGAR, CourtListener, Caselaw Access Project, MTSamples, SROIE, U.S. government forms) and 501 synthetic documents, each labeled by source. A datasheet documents composition, provenance, and per-source licensing. Read before quoting any accuracy number: ~94% of documents are extraction-from-clean-text rather than rendered-page extraction, and synthetic documents overestimate accuracy relative to real ones — always report results stratified by source. See the datasheet in the repository. The scorer and run harness are at https://github.com/fieldbench/fieldbench.



