Wildfire Size Disparity on Tribal Lands in California and Associated PM2.5 Burden, 2000–2018
收藏资源简介:
Version 8.0 — August 2026 Corrected version, published prior to any reviewer feedback. This version self-discloses and corrects three construction errors identified through the authors' own post-submission audit of the v7 replication code. No external reviewer or third party identified these issues; they were found through the authors' internal review process. Full details below. The manuscript (currently under review at Global Environmental Change, GEC-D-26-01575) is being revised in parallel to reflect these corrections. Study Overview This dataset supports a quantitative investigation of wildfire size disparity between federally recognized tribal lands and non-tribal lands in California from 2000 to 2018, and the associated PM2.5 air quality burden borne by tribal communities. Using spatial analysis, panel econometrics, and multiple causal identification strategies, the study documents that fires on tribal trust land are on average 6.4 times larger than fires on non-tribal land — a disparity that persists at 3.1× after excluding boundary-straddling fires and at 4.8× after excluding the catastrophic 2003 and 2007 fire years. All specifications are statistically significant (Mann-Whitney p < 0.0001). v8 Correction Summary Three construction errors were identified and corrected in this version: Tribal boundary selection (data source, Sections 3–4 of the replication code). v7 selected "California" tribal areas using a rectangular latitude/longitude bounding box. Because AIANNH (tribal) boundary polygons in Census TIGER data have no state-code field, this was used as a shortcut — but a rectangle does not respect the actual state line. This included 32 of 185 candidate polygons that are actually in Nevada or Arizona (e.g., Pyramid Lake Paiute, Walker River, Duck Valley, Reno-Sparks, Cocopah). Corrected to use a true geometric intersection with the actual California state boundary, yielding 153 confirmed California tribal areas. This did not change the 111-tribal-fire classification or the 6.4× headline disparity (none of the 111 tribal fires were matched to a non-California polygon), but it materially affected the two geometry-dependent analyses below. Regression discontinuity / propensity score matching running variable (Section 16–17). The RD running variable's sign was assigned from a perimeter-intersection label rather than true point-in-polygon containment of the fire centroid (60% of the 111 tribal fires — 67 of 111 — have centroids that fall outside their own matched tribal polygon despite being labeled tribal by perimeter overlap). Distance was measured to the union of all California tribal land rather than to each fire's own matched parcel, and 51 tribes have multiple non-contiguous, identically-named parcels (5.5–20.6 km apart) that a naive name-based lookup could conflate with the wrong parcel. All three issues were corrected. Fire-to-county assignment (Section 9). v7 joined fires to counties via perimeter intersection, then deduplicated multi-county fires by sorting on total fire acreage — an attribute identical across all of a fire's county-matches, so the sort did not actually break the tie. County assignment for boundary-straddling fires was effectively arbitrary. Corrected to assign each fire to the county with the largest geometric overlap. 126 of 4,468 resolved fires (2.8%), including 5 tribal fires, are assigned to a different county under the fix, changing the tribal-fire indicator for 6 county-year observations. Effect of the corrections: The regression discontinuity result is no longer statistically significant at any bandwidth (see Table below). The original v7 estimates (+394% to +179%, p<0.0001 at all four bandwidths) do not survive correction. The propensity score matching estimate is reduced but remains highly significant: from +2,143% (v7) to a corrected primary estimate of +1,232.3% (caliper-free full match, all 111 tribal fires matched, p=9.4×10⁻¹²). A full caliper sensitivity sweep is included and documented (see below); the point estimate ranges from +1,232% to +4,034% depending on caliper strictness, but statistical significance is robust across every specification tested (p<10⁻¹¹ throughout). The primary two-way fixed-effects PM2.5 estimate is reduced from statistically significant to marginal: from 0.845 µg/m³ (p=0.009) to 0.716 µg/m³ (p=0.064). The event-study pre-trend result is unaffected: no significant pre-trend before or after correction (F=1.72→1.55, both p>0.20). The raw 6.4× fire-size disparity and its sensitivity variants (3.1×, 4.8×, 5.3×) are unaffected by any of the above — they do not depend on the corrected geometry. Given these results, the paper's causal identification case is now better described as one robust descriptive disparity (6.4×) supported by two of four originally reported strategies (PSM, still highly significant; event study, no pre-trend detected) plus one marginal result (primary DiD) and one non-significant result (RD), rather than "four convergent identification strategies" as characterized in the v7 archive description. The authors consider the underlying disparity and its associated PM2.5 health burden to remain well-supported; the RD-specific causal claim does not. Causal Identification Results (Corrected) Strategy 1 — Regression Discontinuity (point-in-polygon sign, matched-parcel distance) Bandwidth Effect p-value Significant? 10 km +18% 0.661 No 25 km +3% 0.943 No 50 km −2% 0.947 No 100 km −6% 0.852 No Strategy 1b — NASA FIRMS Lightning Climatology (reduced form only, no 2SLS) Weak-instrument diagnostics unchanged from v7 (F<10 for both available lightning instruments; 2SLS not reported). Reduced-form estimate not available in this run due to a NASA data-server outage during archive preparation; will be regenerated and added in a subsequent update if source data access is restored. Strategy 2 — Propensity Score Matching (corrected matched-parcel distance covariate) Primary estimate (caliper-free full match, all 111 treated fires): ATT = +1,232.3% (log-acres difference = 2.590), paired t-test p=9.4×10⁻¹², Wilcoxon signed-rank p<0.0001. Caliper sensitivity (documented transparently rather than reporting a single point estimate): Caliper Matched pairs ATT p-value 0.01 78 +4,034% 7.8×10⁻¹⁶ 0.02 83 +3,890% 1.8×10⁻¹⁶ 0.03 88 +3,204% 4.2×10⁻¹⁶ 0.05 93 +2,665% 1.0×10⁻¹⁵ 0.08 97 +2,158% 1.1×10⁻¹⁴ 0.10 103 +1,758% 1.1×10⁻¹³ 0.15 110 +1,259% 9.7×10⁻¹² 0.20–0.30 (full match) 111 +1,232% 9.4×10⁻¹² Magnitude is sensitive to caliper choice (tighter calipers select a subset of closer-propensity matches with larger size gaps); statistical significance is not (p<10⁻¹¹ at every caliper tested, and under match-order permutation). Covariate balance after matching remains imperfect (standardized mean differences of 0.34–0.37 on latitude, longitude, and distance-to-tribal-land), noted here as a limitation. Strategy 3 — Event Study Pre-Trend Test No significant pre-trend: F=1.549, p=0.246 (t=−5 to −2, relative to first tribal fire year per county). Significant post-period effect: F=6.044, p=0.0033. Primary DiD — Two-Way Fixed Effects OLS (county + year FE, clustered SE) Tribal fire coefficient: 0.716 µg/m³, p=0.0635 (95% CI: −0.040 to 1.471). Not significant at the conventional 0.05 threshold; marginal. Boundary Stability Tests (Section 16d — corrected) Test 1 — Boundary vintage check: all 153 corrected California tribal boundaries carry federally recognized MTFCC designations (G2101/G2102/G2160). Unchanged conclusion from v7: boundary endogeneity ruled out by legal structure. Test 2 — Placebo boundary RD (label permutation, 500 permutations, 25 km bandwidth): re-run on the corrected running variable. The real effect (+2.7%, p=0.943) is not extreme relative to the placebo distribution (permutation p≈0.97, roughly 483–489 of 500 placebos exceed the real estimate across repeated runs) — consistent with, and independently corroborating, the corrected RD's null result. This reverses the v7 finding (permutation p=0.006, real estimate an extreme outlier), which was itself an artifact of the same running-variable construction error described above. Test 3 — Covariate balance at boundary: centroid latitude and longitude remain significantly discontinuous at the boundary (p<0.001), reflecting the geographic concentration of California reservations in distinct ecological zones, as in v7. Log fire size (the outcome) is no longer significantly discontinuous (p=0.943), consistent with the corrected null RD result rather than the "expected discontinuity" reported in v7. Fire Size Disparity (unaffected by geometry corrections) Specification Tribal mean (acres) Non-tribal mean (acres) Ratio p-value Full dataset (deduplicated) 17,265 2,706 6.4× <0.0001 Excl. 2003 14,845 2,782 5.3× <0.0001 Excl. 2003 + 2007 13,545 2,808 4.8× <0.0001 Excl. boundary-straddlers 8,253 2,706 3.1× <0.0001 Managing Agency Distribution (Tribal Trust Lands) The Bureau of Indian Affairs — the only federal agency with statutory trust responsibility toward tribal nations — managed 9.0% of fires on tribal trust land during the study period (10 of 111 deduplicated tribal fires), not the 7.0% reported in the v7 archive description, which was computed on a pre-deduplication 143-fire count. The full managing-agency breakdown (CAL FIRE, state agencies, USFS, BLM, other) is being regenerated against the corrected 111-fire dataset and will be added to this archive in a subsequent file update; only the BIA figure has been independently re-verified as of this description. Data Contents Files included (19 files): FIRE_v8.ipynb — Full integrated analysis pipeline (Sections 1–22), corrected as described above. Contains the complete revision notes, methodology, and all analysis code. README.md — Archive overview, file descriptions, correction summary, and reproduction instructions ca_fires_tribal_classified.csv — 4,496 California fire perimeters with tribal land classification (111 tribal after deduplication) tribal_fire_pm25_panel.csv — County-year observations (PM2.5 + fire variables), rebuilt on the corrected largest-overlap county assignment ca_county_year_pm25.csv — County-year PM2.5 summary statistics from EPA AQS tribal_nation_fire_burden.csv — Fire burden summary by tribal nation event_study_coefficients.csv — Event study coefficients (t = −5 to +5) with 95% CI, corrected panel regression_results.txt — Full two-way FE OLS output including all fixed effects, corrected boundary_stability_table.csv — Section 16d covariate-balance and placebo test results, corrected (machine-readable) boundary_stability_results.txt — Section 16d results, corrected (narrative summary; generated with a fixed random seed for reproducibility — see file header for details on seed-to-seed variation) placebo_rd_distribution.csv — 500 permutation placebo RD estimates, corrected psm_caliper_sensitivity.csv — NEW in v8: full caliper sensitivity sweep for the PSM estimate (see table above) figure1_tribal_fire_analysis.png through figure7_causal_identification.png (7 files) — Publication figures, regenerated from corrected pipeline This archive contains data files, the analysis notebook, and documentation; there is no separate replication .zip bundle, as the notebook and its outputs together are fully sufficient to reproduce all results. Methods Summary Fire perimeter data were obtained from the NIFC Historical GeoMAC Perimeters archive (2000–2018) via ArcGIS REST API. Tribal boundaries are from the Census TIGER/Line AIANNH 2025 vintage, filtered to a true California state-boundary intersection (153 confirmed California areas, corrected from a 185-polygon bounding-box selection in v7). PM2.5 concentrations are from EPA AQS daily FRM/FEM monitoring data (parameter 88101; 580,066 California daily observations). All spatial operations use the California Albers Equal Area projection (EPSG:3310) for area, centroid, and distance calculations. Spatial join deduplication removed 12 NIFC duplicate polygon records and resolved 8 boundary-straddling fires to their largest-intersection tribal match, yielding 111 deduplicated tribal fires from 143 raw matches (unchanged from v7). Causal identification employs four strategies, detailed above: (1) regression discontinuity at tribal land boundaries (four bandwidth specifications, corrected geometry); (2) propensity score matching (corrected distance covariate, caliper-free primary specification with full sensitivity sweep); (3) event study with county and year fixed effects and pre-trend joint F-test; (4) two-way fixed effects DiD (county + year FE, clustered SE, corrected county assignment). Revision History v8.0 (August 2026): Self-audited correction, published prior to any external reviewer feedback. Corrects three construction errors: (1) tribal boundary selection used a bounding box rather than a true state-boundary intersection, including 32 non-California polygons; (2) RD/PSM running variable used a perimeter-based sign label and distance-to-union rather than true point-in-polygon containment and matched-parcel distance; (3) fire-to-county assignment used a non-discriminating tie-break rather than largest-overlap assignment. Effects: RD no longer significant at any bandwidth; PSM ATT reduced from +2,143% to +1,232% (still highly significant, full caliper sensitivity now documented); primary DiD reduced from 0.845 µg/m³ (p=0.009) to 0.716 µg/m³ (p=0.064); event study conclusion unchanged. BIA managing-agency figure corrected from 7.0% to 9.0% (deduplicated basis). Full pipeline independently re-verified across three consecutive end-to-end test runs prior to publication of this version. v7.0 (June 2026): Integrates boundary stability tests (Section 16d) into canonical deduplicated pipeline. Three tests conducted: TIGER vintage check, label permutation placebo RD, and covariate balance at boundary. All tests run on correct 111-fire deduplicated dataset. Superseded by v8 as described above. v6.0 (June 2026): Boundary stability outputs generated from pre-deduplication dataset — superseded by v7. v5.0 (June 2026): Zenodo archive with complete replication pipeline. v2.0 (June 2026): Critical methodological corrections — spatial join deduplication (143 → 111 tribal fires), two-way fixed effects added, 2SLS removed (weak instruments), boundary-straddler sensitivity documented. v1.0 (June 2026): Initial deposit. Computational Environment Python 3.12.13 · geopandas 1.1.4 · pandas 2.2.2 · statsmodels 0.14.6 · scipy 1.16.3 · scikit-learn 1.6.1 · netCDF4 1.7.4 · Executed on Google Colab CPU runtime · Approximate runtime: 30–40 minutes (includes live download from NIFC, Census TIGER, and EPA AQS) License Creative Commons Attribution 4.0 International (CC BY 4.0). Users are free to share and adapt the material for any purpose provided appropriate credit is given to the authors.



