遇见数据集

FieldBench Corpus: A Cross-Domain Benchmark for Schema-Driven Document Extraction

收藏
Zenodo2026-07-24 更新2026-08-02 收录
官方服务:

资源简介:

A cross-domain, field-level benchmark for schema-driven document extraction (document → structured JSON): 1,114 documents across 12 categories with per-field ground truth, released so extraction-accuracy claims become falsifiable and comparable. 613 real documents (SEC EDGAR, CourtListener, Caselaw Access Project, MTSamples, SROIE, U.S. government forms) and 501 synthetic documents, each labeled by source. A datasheet documents composition, provenance, and per-source licensing. Read before quoting any accuracy number: ~94% of documents are extraction-from-clean-text rather than rendered-page extraction, and synthetic documents overestimate accuracy relative to real ones — always report results stratified by source. See the datasheet in the repository. The scorer and run harness are at https://github.com/fieldbench/fieldbench.

提供机构:
Zenodo
创建时间:
2026-07-24
二维码
社区交流群
二维码
科研交流群
商业服务