Documented Need: How Federal Lead Pipe Funding Is Allocated on the Quality of Water Systems' Paperwork — replication package (code, checksummed raw archive, processed panel)
收藏资源简介:
Replication package for the paper "Documented Need: How Federal Lead Pipe Funding Is Allocated on the Quality of Water Systems' Paperwork" (Gustavo Pedro Ricou, Trinity College Dublin). The paper reverse-engineers the EPA DWSRF Lead Service Line Replacement allotment formula, documents how the formula converts water systems' documentation practices into money, and estimates the federal return to resolving unknown service lines. Why the manuscript is archived here. The manuscript is archived alongside the code and data not as a second publication venue but because, in this work, the three are one object. Every quantitative claim in the paper is recomputed from the committed measurement data by the archived analysis code, and the archived test suite fails if the manuscript and the data ever disagree — so the version of record of the evidence necessarily includes the version of the text it supports. Depositing them together, under one DOI, in a byte-exact archive of a single tagged commit (v1.0.0), means a reader can verify any sentence of the paper against the data behind it without trusting anything but arithmetic; it also gives the manuscript an independent, third-party timestamp from the moment of release and mirrors the whole object into Software Heritage for long-term preservation. A paper about a formula that pays on unverified paperwork should not itself ask for trust it can decline to need: this record is that principle applied to the paper itself. Contents manuscript-documented-need-v1.0.pdf — the manuscript as compiled from the tagged sources, citing this record's DOI. lead-hidden-need-code-v1.0.zip — full source tree at tagged commit v1.0.0: analysis pipeline (Python 3.9; pandas, numpy, statsmodels), 860 automated tests at 100% branch coverage, documentation registers, manuscript sources, and per-file checksummed provenance sidecars (*.meta.json) for the raw archive. Code is MIT licensed (LICENSE inside). data-raw-inventory-v1.0.zip — six quarterly vintages of EPA's service line inventory dashboard exports (2025Q1–2026Q2). Five of the six no longer exist upstream: EPA overwrites the dashboard each quarter, so this panel cannot be reconstructed from public sources today. data-raw-federal-other-v1.0.zip — EPA allotment memoranda, 7th DWINSA materials, eCFR / Federal Register / U.S. Code policy vintages, Census and RUCA extracts, DWSRF materials. data-raw-states-grey-utilities-v1.0.zip — state program materials, utility-level validation data (Illinois, Michigan, Norfolk, New York), grey literature, manual acquisitions, and the quarantine folder. data-raw-federal-echo-v1.0.zip — EPA ECHO compliance-history extracts. data-raw-federal-boundaries-v1.0.zip — EPA community water system service-area boundary data. data-processed-v1.0.zip — the processed analysis panel (parquet) and reference tables, allowing replication without rebuilding from raw. Integrity. Every raw file carries a .meta.json sidecar recording its source URL, acquisition date and SHA-256 digest; the code archive includes the verification tool (python -m lead_hidden_need.acquire.verify_archive), and SHA256SUMS.txt lists the digest of every file in this record. All archived files verify against their recorded digests at deposit. Licensing. The record license (CC BY 4.0) applies to the data compilation and its documentation. The code is MIT licensed (LICENSE in the code archive). The underlying federal datasets are U.S. Government works. The manuscript file (manuscript-documented-need-v1.0.pdf) is © 2026 Gustavo Pedro Ricou, included for verifiability and archival timestamping, and is not licensed under CC BY pending journal publication. Code repository: https://github.com/Gustolandia/lead-hidden-need (tag v1.0.0). A preprint of the manuscript has been submitted to arXiv; the identifier will be added to this record when announced.



