Replication data for \"Evaluating Bias and Noise Induced by the U.S. Census Bureau's Privacy Protection Methods\"
收藏资源简介:
The United States Census Bureau faces a difficult trade-off between the accuracy of Census statistics and the protection of individual information. We conduct the first independent evaluation of bias and noise induced by the Bureau's two main disclosure avoidance systems: the TopDown algorithm employed for the 2020 Census and the swapping algorithm implemented for the three previous Censuses. Our evaluation leverages the Noisy Measure File (NMF) as well as two independent runs of the TopDown algorithm applied to the 2010 decennial Census. We find that the NMF contains too much noise to be directly useful, especially for Hispanic and multiracial populations. TopDown's post-processing dramatically reduces the NMF noise and produces data whose accuracy is similar to that of swapping. While the estimated errors for both TopDown and swapping algorithms are generally no greater than other sources of Census error, they can be relatively substantial for geographies with small total populations.
美国人口普查局在人口普查统计数据的准确性与个人信息保护之间面临艰难权衡。本研究首次对该局两大主要披露规避系统所引发的偏差与噪声开展独立评估:分别为2020年人口普查所用的TopDown算法,以及此前三次人口普查采用的交换算法。本次评估依托噪声测量文件(Noisy Measure File,缩写NMF),以及针对2010年十年一次人口普查开展的两次独立TopDown算法运行结果。研究发现,NMF所含噪声过大,无法直接投入使用,针对西班牙裔及多种族群体的情况尤为如此。TopDown算法的后处理流程可大幅降低NMF的噪声水平,生成的数据精度与交换算法相当。尽管TopDown算法与交换算法的估算误差总体上不高于人口普查的其他误差来源,但在总人口规模较小的地理区域中,这类误差相对较为显著。



