Differentially Private Methods for Compositional Data
收藏资源简介:
Confidential data, such as electronic health records, activity data from wearable devices, and geolocation data, are becoming increasingly prevalent. Differential privacy provides a framework to conduct statistical analyses while mitigating the risk of leaking private information. Compositional data, which consist of vectors with positive components that add up to a constant, have received little attention in the differential privacy literature. This article proposes differentially private approaches for analyzing compositional data based on the Dirichlet distribution. We explore several methods, including Bayesian and bootstrap procedures. For the Bayesian methods, we consider posterior inference techniques based on Markov chain Monte Carlo, Approximate Bayesian Computation, and asymptotic approximations. We conduct an extensive simulation study to compare these approaches and make evidence-based recommendations. Finally, we apply the methodology to a dataset from the American Time Use Survey.
诸如电子健康记录(electronic health records)、可穿戴设备活动数据以及地理位置数据在内的机密数据正日益普及。差分隐私(differential privacy)提供了一种可在开展统计分析的同时降低隐私信息泄露风险的研究框架。成分数据(compositional data)指由各分量均为正数且总和为定值的向量构成的数据,目前在差分隐私相关研究领域中尚未得到足够关注。本文提出了基于狄利克雷分布(Dirichlet distribution)的差分隐私分析方法,用于成分数据的统计分析。本文探讨了多种分析路径,涵盖贝叶斯(Bayesian)方法与自助法(bootstrap procedures);针对贝叶斯方法,本文研究了基于马尔可夫链蒙特卡洛(Markov chain Monte Carlo, MCMC)、近似贝叶斯计算(Approximate Bayesian Computation, ABC)以及渐近近似的后验推断方法。本文开展了大规模模拟研究以对比上述各类方法,并给出循证推荐方案。最后,本文将所提方法论应用于美国时间使用调查(American Time Use Survey)的数据集。



