Applying Data Synthesis for Longitudinal Business Data across Three Countries
收藏资源简介:
Data on businesses collected by statistical agencies are challenging to protect.Many businesses have unique characteristics, and distributions of employment,sales, and profits are highly skewed. Attackers wishing to conduct identificationattacks often have access to much more information than for any individual. Asa consequence, most disclosure avoidance mechanisms fail to strike an accept-able balance between usefulness and confidentiality protection. Detailed aggregatestatistics by geography or detailed industry classes are rare, public-use microdataon businesses are virtually inexistant, and access to confidential microdata can beburdensome. Synthetic microdata have been proposed as a secure mechanism topublish microdata, as part of a broader discussion of how to provide broader accessto such datasets to researchers. In this article, we document an experiment to cre-ate analytically valid synthetic data, using the exact same model and methods previ-ously employed for the United States, for data from two different countries: Canada(Longitudinal Employment Analysis Program (LEAP)) and Germany (EstablishmentHistory Panel (BHP)). We assess utility and protection, and provide an assessmentof the feasibility of extending such an approach in a cost-effective way to other data.
统计机构采集的企业相关数据,其保护工作极具挑战性。多数企业具备独特属性,就业人数、销售额与利润的分布呈现高度偏态特征。意图开展身份识别攻击的攻击者,往往可获取远超普通个体的信息。因此,多数披露规避机制难以在实用性与机密性保护之间达成可接受的平衡。按地理区域或细分行业类别划分的详细汇总统计数据极为稀缺,公开可用的企业微数据几乎不存在,而访问机密微数据的流程往往颇为繁琐。作为“为研究人员提供此类数据集更广泛访问途径”这一整体讨论的组成部分,合成微数据(synthetic microdata)已被提出作为一种安全的微数据发布机制。本文中,我们开展了一项实验:针对来自加拿大的纵向就业分析项目(Longitudinal Employment Analysis Program,LEAP)与德国的企业历史面板(Establishment History Panel,BHP)数据集,采用此前用于美国的完全一致的模型与方法,生成具备分析有效性的合成数据。我们对该合成数据的效用与防护性能进行评估,并探讨了以成本效益高的方式将此类方法推广至其他数据集的可行性。




