Medicaid Section 1115 Demonstration Evaluation Findings: an open dataset (Dark Health Data)
收藏资源简介:
Dark Health Data is an open-source project that turns buried public-record health documents (PDFs) into research-ready, provenance-stamped datasets via large language model extraction (Anthropic Claude claude-haiku-4-5) with a verification layer: source grounding, symbolic and neurosymbolic (Logical Neural Network-inspired) constraint checks, a heterogeneous model ensemble, and a conformal selective-acceptance gate. Source documents are independent evaluations of Medicaid Section 1115 demonstration waivers (including health-related social needs pilots for food and housing), which report evaluation findings and recommendations. This release (v0.3.0) contains 7,198 records extracted from 12 source documents across 7 states/jurisdictions: 4,525 evaluation findings, 2,673 recommendations. Every record carries full provenance (source-document SHA-256, page, extraction model, and confidence) and a quality/trust score; nothing is imputed or dropped — problems are flagged with QA codes and a low trust score. These records are AI-extracted and not yet independently validated — a preliminary first release with no accompanying validation study. Treat the data as preliminary, filter on the per-record trust score, and verify values against the linked source pages. Derived entirely from public-record documents; contains no protected health information and is not human-subjects research. Code (Apache-2.0): https://github.com/sanjaybasu/dark-health-data. Data license: CC0-1.0.



