Hospital Community Health Needs Assessment (CHNA) Findings: an open dataset (Dark Health Data)
收藏资源简介:
Dark Health Data turns buried public-record health documents (PDFs) into research-ready, provenance-stamped datasets via large language model extraction (Anthropic Claude claude-haiku-4-5) with a verification layer (grounding, neurosymbolic constraints, ensemble, conformal gate). Source documents are U.S. non-profit hospital Community Health Needs Assessments (CHNAs), required triennially under IRC 501(r)(3). This release (v0.4.2) is an expanded national crawl: 473,319 records from 1,394 source documents naming 5,039 hospitals across 51 states (2-letter normalized): 366,629 identified community health needs, 106,690 implementation strategies. The crawl is ongoing toward the full ~2,900 non-profit-hospital universe; later versions will add more facilities. Every record carries full provenance and a quality/trust score; nothing is imputed or dropped. AI-extracted, not yet independently validated — preliminary; filter on the trust score and verify against the linked source pages. Public-record documents only; no PHI. Code (Apache-2.0): github.com/sanjaybasu/dark-health-data. CC0-1.0.



