遇见数据集

SIY Longitudinal Health-Statistics Extraction Framework: Database, Pipeline and Discrepancy Catalogue

收藏
Zenodo2026-07-13 更新2026-06-05 收录
官方服务:

资源简介:

Open dataset and Python pipeline accompanying a longitudinal table-extraction study on the seventeen-volume Turkish Ministry of Health Statistics Yearbook (Sağlık İstatistikleri Yıllığı, SIY, 2008–2024). This deposit provides:- a 69 MB unified SQLite database integrating four data layers: extraction output, curated indicator catalogue, inter-edition data-quality discrepancies, and a full-text page corpus;- a 27-script Python pipeline that reproduces every step of the cascade (rasterisation → Table-Transformer detection → structure recovery → LayoutLMv3 multimodal cell understanding → ByT5 byte-level encoding repair → schema normalisation);- all six figures of the accompanying manuscript and a graphical abstract;- an interactive analytics dashboard (SIY_Dashboard_Pro.html) referenced in Fig. 5 of the manuscript. Released ground truth: 2,651 tables, 268,109 cells, across 17 editions.Curated layer: 46 indicators with 2,186 long-format observations.Data-quality layer: 63 documented inter-edition value discrepancies (34 major + 29 minor).Full-text page corpus: 7,725 pages (4,083 TR + 3,642 EN). Evaluation on the held-out 2024 test split (166 tables, 19,980 cells): TEDS-Struct 0.91, TEDS-Content 0.90, Cell-F1 0.87, a 22-fold Cell-F1 improvement over pdfplumber baseline. License: data and figures under CC BY-SA 4.0; Python code under MIT. This deposit accompanies the manuscript "A Longitudinal Table Extraction Framework for Evidence-Based Health Policy" (PeerJ Computer Science, AI Application track, Manuscript ID CS-2026:05:141119). CONTENTS1. data/SIY_Veritabani_birlesik.db: Consolidated SQLite database (69 MB) containing seven relations: yearbooks (17 editions), tables (2,651), cells (268,109), indicators (46), values_long (2,186), discrepancies (63), and pdf_pages (7,725 pages).2. code/: Twenty-seven Python scripts, split into six extraction scripts (corpus acquisition and raw table extraction), twenty curation scripts (indicator catalogue construction, inter-edition reconciliation, and discrepancy detection), and one baseline-evaluation script (code/evaluation/run_docling_baseline.py that reproduces the Docling row of Table 4).3. SIY_Dashboard_Pro.html: Interactive analytics dashboard referenced in Fig. 5.4. figures/: All six manuscript figures plus a graphical abstract.5. README.md and README.docx: Reproduction guide, including the modern document-AI baseline comparison protocol.6. CITATION.docx: Recommended citation and BibTeX entry.7. LICENSE-code.docx: MIT License covering all Python code.8. LICENSE-data.docx: Creative Commons Attribution-ShareAlike 4.0 International License covering the SQLite database and derived data. VERSION HISTORY- v1.5.0 (13 July 2026): Added SIY_Dashboard_Pro.html; reconciled the reported pipeline script count with the actual archive content (27 scripts, previously stated as 25); translated remaining Turkish inline comments to English; refined the Acknowledgements and AI-use disclosure.- v1.4.0 (13 July 2026): Added SIY_Dashboard_Pro.html; corrected script count in README; removed FUSE and WAL artefacts.- v1.2.0 (17 June 2026): Added code/evaluation/run_docling_baseline.py and revised README to document the modern document-AI baseline comparison requested by reviewers.- v1.1.0 (16 June 2026): Initial public release accompanying the PeerJ submission. REPRODUCIBILITYThe 17-yearbook source PDF corpus (approximately 1.2 GB) is publicly redistributed by the Republic of Türkiye Ministry of Health and can be re-downloaded via code/00_download_siy_pdfs.py. The deposit is self-contained: a downstream analyst can re-run the full pipeline on the released PDFs, verify the 2,651 tables and 268,109 cells, reproduce the 63-entry discrepancy catalogue, and independently benchmark Docling on the same evaluation slice.

提供机构:
Zenodo
创建时间:
2026-05-29
二维码
社区交流群
二维码
科研交流群
商业服务