遇见数据集

SIY Longitudinal Health-Statistics Extraction Framework: Database, Pipeline and Discrepancy Catalogue

收藏
Zenodo2026-05-29 更新2026-06-05 收录
官方服务:

资源简介:

Open dataset and Python pipeline accompanying a longitudinal table-extraction study on the seventeen-volume Turkish Ministry of Health Statistics Yearbook (Sağlık İstatistikleri Yıllığı, SIY, 2008–2024).This deposit provides: • a 69 MB unified SQLite database integrating four data layers — extraction output, curated indicator catalogue, inter-edition data-quality discrepancies, and a full-text page corpus; • a 24-script Python pipeline that reproduces every step of the cascade (rasterisation → Table-Transformer detection → structure recovery → LayoutLMv3 multimodal cell understanding → ByT5 byte-level encoding repair → schema normalisation); • all six figures of the accompanying manuscript and a graphical abstract; • a manuscript preprint (.docx).Released ground truth: 2,651 tables, 268,109 cells, across 17 editions.Curated layer: 46 indicators with 2,186 long-format observations.Data-quality layer: 63 documented inter-edition value discrepancies (34 major + 29 minor).Full-text page corpus: 7,725 pages (4,083 TR + 3,642 EN). Evaluation on the held-out 2024 test split (166 tables, 19,980 cells): TEDS-Struct 0.91, TEDS-Content 0.90, Cell-F1 0.87 — a 22-fold Cell-F1 improvement over pdfplumber baseline.License: data and figures under CC BY-SA 4.0; Python code under MIT.

提供机构:
Zenodo
创建时间:
2026-05-29
二维码
社区交流群
二维码
科研交流群
商业服务