Replication data for: Fraudulent Democracy? An Analysis of Argentina's Infamous Decade using Supervised Machine Learning
收藏资源简介:
In this paper we introduce an innovative method to diagnose electoral fraud using vote counts. Specically, we use synthetic data to develop and train a fraud detection prototype. We employ naive Bayes classier as our learning algorithm and rely on digital analysis to identify the features that are most informative about class distinctions. To evaluate the detection capability of the classier we use authentic data drawn from a novel dataset of district-level vote counts in the province of Buenos Aires (Argentina) between 1931 and 1941, a period with a checkered history of fraud. Our results corroborate the validity of our approach: the elections considered to be irregular (legitimate) by most historical accounts are unambiguously classied as fraudulent (clean) by the learner. More generally, our ndings demonstrate the feasibility of generating and using synthetic data for training and testing an electoral fraud detection system.



