ML-ARDD: Simulation Code and Data for Machine Learning-Based Adaptive Bandwidth Selection in Regression Discontinuity Designs
收藏资源简介:
This repository contains the full simulation code and synthetic dataset supporting the manuscript "A Machine Learning-Based Adaptive Framework for Bandwidth Selection in Regression Discontinuity Designs: Simulation Evidence." It includes the R Markdown source implementing ML-ARDD (Generalised Random Forests, Gradient Boosting Machines, and Support Vector Machines for empirical bias–variance-based bandwidth selection), the Monte Carlo simulation pipeline (43,200 simulation rows, 500 replications per cell, two data-generating process scenarios), and all sensitivity and robustness analyses reported in the manuscript. The accompanying dataset (simulated_health_data.csv, N = 5,000) is a fully synthetic dataset used in the single-pass applied study; no individual-level real-world records are included or required to reproduce any result in this work.



