S 4320 Project 1: Predicting Competitive US House District Outcomes Using Demographic and Structural Features
收藏资源简介:
A longitudinal secondary dataset covering all 435 US House of Representatives congressional districts across six election cycles from 2012 through 2022. The dataset combines certified electoral returns from the MIT Election Data and Science Lab, district-level demographic estimates from the US Census Bureau American Community Survey 5-year estimates, and manually compiled national economic and political context variables including generic ballot margins, presidential approval ratings, GDP growth, and unemployment rates. The dataset is structured as five relational tables: a districts dimension table, an elections fact table with vote outcomes and incumbency flags, a demographics fact table with 18 socioeconomic features per district per cycle, a national context table with macro-level predictors by year, and a results table with gradient boosting model predictions for party seat flips. All tables are stored in parquet format and linked by shared district and year keys. Key features include lagged electoral margins, party flip indicators, Democratic vote swing, competitiveness scores, urban-rural classification, poverty rate, health insurance coverage rate, median age, and median home value. A gradient boosting classifier trained on the dataset predicts which districts will flip party with a ROC-AUC of 0.98 on a held-out 2022 test set using no polling data. Intended for researchers studying electoral forecasting, demographic change and partisan realignment, and the structural determinants of congressional competitiveness.



