Hospital Community Health Needs Assessment (CHNA) Findings: an open dataset (Dark Health Data)
收藏资源简介:
Dark Health Data is an open-source project that turns buried public-record health documents (PDFs) into research-ready, provenance-stamped datasets via large language model extraction (Anthropic Claude claude-haiku-4-5) with a verification layer: source grounding, symbolic and neurosymbolic (Logical Neural Network-inspired) constraint checks, a heterogeneous model ensemble, and a conformal selective-acceptance gate. Source documents are U.S. non-profit hospital Community Health Needs Assessments (CHNAs), required triennially under IRC §501(r)(3), which identify and prioritize community health needs and set implementation strategies. This release (v0.3.0) contains 9,152 records extracted from 27 source documents across 96 hospitals: 6,960 identified community health needs, 2,192 implementation strategies. Every record carries full provenance (source-document SHA-256, page, extraction model, and confidence) and a quality/trust score; nothing is imputed or dropped — problems are flagged with QA codes and a low trust score. These records are AI-extracted and not yet independently validated — a preliminary first release with no accompanying validation study. Treat the data as preliminary, filter on the per-record trust score, and verify values against the linked source pages. Derived entirely from public-record documents; contains no protected health information and is not human-subjects research. Code (Apache-2.0): https://github.com/sanjaybasu/dark-health-data. Data license: CC0-1.0.



