Raw fitted models and Shapley intermediates for Replication code
收藏资源简介:
This record holds the heavy inputs of the replication package for Borup, Goulet Coulombe, Rapach, Montes Schütte and Schwenk-Nebbe, "The Anatomy of Out-of-Sample Forecasting Accuracy: A Shapley-Based Approach", Journal of Applied Econometrics (forthcoming). The replication package itself, with the code, the input data, the saved estimation artifacts, and the scripts that regenerate every table and figure of the paper and its Internet Appendix, is available at https://github.com/Christianmontes/anatomy_of_out_of_sample_forecast_accuracy and in the JAE Data Archive. That package is self-sufficient for all published results. The files in this record are needed only to recompute the saved Shapley and GPBSV artifacts from the raw fitted models (Tier 2 in the package README) or to refit the forecasting models (Tier 3). Contents, 15 files, about 23 GB: 20241005_091334.7z (7.2 GB compressed, 27.4 GB unpacked): the fitted forecasting models for US CPI inflation at horizons 1, 3, 6 and 12 months, 416 rolling-window estimation periods per horizon (out-of-sample period 1990:01 to 2024:08), run id 20241005_091334, random seed 43210. Models: principal component regression, elastic net, random forest, XGBoost, two neural networks and the AR benchmark. anatomy_h1_upd.bin, anatomy_h3_upd.bin, anatomy_h6_upd.bin, anatomy_h12_upd.bin (about 1.2 GB each): the raw GPBSV permutation draws produced by iml_rev.estimate(h). ishapley_h1_upd.bin, ishapley_h3_upd.bin, ishapley_h6_upd.bin, ishapley_h12_upd.bin (1.6 to 1.8 GB each): the in-sample Shapley value arrays produced by iml_rev_ishapley.estimate(h) and consumed by iml_rev.anatomize(h). _window_cache_h1.zip, _window_cache_h3.zip, _window_cache_h6.zip, _window_cache_h12.zip (about 1 GB each): per-window resume caches of the in-sample GPBSV computation behind Internet Appendix Table A.3 (416 files per horizon). Optional; the aggregated results ship with the package. README.md and SHA256SUMS.txt: the generating script and the restore location of each file, and SHA-256 checksums (verify with sha256sum -c SHA256SUMS.txt). Usage: place each file at the location given in the package's Models/README.md (the 7z is extracted into Models/20241005_091334/, the bin files go into Results/Updated CPI 3/, the zips are extracted into revision/mas_referee_response/outputs/insample_gpbsv_v5/). The code reads the model folder from the environment variable IML_MODELS_DIR. Tier 2 needs Python 3.9 with the package's pinned environment plus tensorflow 2.5.0 and shap 0.42.1. Data behind the models: FRED-MD, vintage of 5 October 2024 (McCracken and Ng, 2016, Federal Reserve Bank of St. Louis), and the Index of Consumer Sentiment, Index of Consumer Expectations and Index of Current Economic Conditions of the University of Michigan Surveys of Consumers. The input data are part of the replication package, not of this record. Please cite the paper when using these files.



