Replication Data for: How Cross-Validation Can Go Wrong and What to Do About it.
收藏资源简介:
The introduction of new “machine learning” methods and terminology to political science complicates the interpretation of results. Even more so, when one term – like cross-validation – can mean very different things. We find different meanings of cross-validation in applied political science work. In the context of predictive modeling, cross-validation can be used to obtain an estimate of true error or as a procedure for model tuning. Using a single cross-validation procedure to obtain an estimate of the true error and for model tuning at the same time leads to serious misreporting of performance measures. We demonstrate the severe consequences of this problem with a series of experiments. We also observe this problematic usage of cross-validation in applied research. We look at Muchlinski et al. (2016) on the prediction of civil war onsets to illustrate how the problematic cross-validation can affect applied work. Applying cross-validation correctly, we are unable to reproduce their findings. We encourage researchers in predictive modeling to be especially mindful when applying cross-validation.
将新兴的机器学习(machine learning)方法与术语引入政治学领域,会加大研究结果解读的难度。更有甚者,诸如交叉验证(cross-validation)这类术语,其含义可能存在显著差异。我们在应用政治学研究中发现,交叉验证存在多种不同的释义。在预测建模场景下,交叉验证可用于获取真实误差的估计值,或是作为模型调优的流程手段。若同时使用单一交叉验证流程来获取真实误差估计值并完成模型调优,则会导致性能指标的严重误报。我们通过一系列实验,展示了该问题带来的严重后果。我们还在应用类研究中观察到了这类存在问题的交叉验证使用方式。我们以穆奇林斯基等人(Muchlinski et al., 2016)关于内战爆发预测的研究为例,阐释存在问题的交叉验证使用方式会如何影响应用研究成果。采用正确的交叉验证方法后,我们无法复现其研究结论。我们呼吁预测建模领域的研究者在使用交叉验证时保持格外谨慎。



