Hospital Community Health Needs Assessment (CHNA) Findings: an open dataset (Dark Health Data)
收藏资源简介:
Dark Health Data turns buried public-record health documents (PDFs) into research-ready, provenance-stamped datasets via large language model extraction (Anthropic Claude claude-haiku-4-5) with a verification layer: source grounding, symbolic + Logical Neural Network-inspired neurosymbolic constraints, a heterogeneous ensemble, and a conformal selective-acceptance gate. Source documents are U.S. non-profit hospital Community Health Needs Assessments (CHNAs), required triennially under IRC 501(r)(3), which identify and prioritize community health needs and set implementation strategies. This release (v0.4.1) contains 277,273 records from 694 source documents naming 3,410 hospitals across 51 states (a system CHNA often names several member hospitals): 207,771 identified community health needs, 69,502 implementation strategies. Every record carries full provenance (source-document SHA-256, page, model, confidence) and a quality/trust score; nothing is imputed or dropped. v0.4.1 normalizes the state field to 2-letter USPS codes (v0.4.0 stored a mix of full names and codes), for consistency with the other Dark Health Data datasets. These records are AI-extracted and not yet independently validated — treat as preliminary, filter on the trust score, and verify values against the linked source pages. Derived entirely from public-record documents; no protected health information. Code (Apache-2.0): github.com/sanjaybasu/dark-health-data. Data license: CC0-1.0.



