Improving the Utility of Poisson-Distributed, Differentially Private Synthetic Data Via Prior Predictive Truncation with an Application to CDC WONDER

Improving the Utility of Poisson-Distributed, Differentially Private Synthetic Data Via Prior Predictive Truncation with an Application to CDC WONDER
复制标题

通过在 CDC WONDER 中的应用,通过先验预测截断提高泊松分布、差分隐私合成数据的实用性

DOI:
10.1093/jssam/smac007
复制
发表时间:
2022
影响因子:
2.1
通讯作者:
Quick, Harrison
Quick, Harrison
中科院分区:
数学3区
文献类型:
--
作者:
Quick, Harrison

文献摘要

相似文献

CDC WONDER是一个基于网络的工具,用于传播国家生命统计系统收集的流行病学数据。虽然CDC WONDER具有内置的隐私保护,但它们不满足正式的隐私保护,例如差分隐私,因此容易受到有针对性的攻击。鉴于高质量的公共卫生数据公开的重要性,同时保护底层数据主体的隐私,我们的目标是提高最近开发的方法的效用,通过使用公开的信息来截断合成数据的范围,生成泊松分布的,不同的私人合成数据。具体来说,我们利用美国人口普查局的县级人口信息和疾病预防控制中心编制的国家死亡报告,为县级死亡率的先验分布提供信息,并推断泊松分布的县级死亡计数的合理范围。在这样做时,对于给定的隐私预算,满足差分隐私的要求可以减少几个数量级,从而导致实用性的实质性改进。为了说明我们提出的方法,我们考虑了一个数据集,该数据集由来自宾夕法尼亚州联邦的超过26,000例癌症相关死亡组成,属于死亡原因和人口统计变量(如年龄,种族,性别和居住县)的47,000多种组合,并证明了所提出的框架能够保留地理,城市/农村,和种族差异的真实数据。
CDC WONDER is a web-based tool for the dissemination of epidemiologic data collected by the National Vital Statistics System. While CDC WONDER has built-in privacy protections, they do not satisfy formal privacy protections such as differential privacy and thus are susceptible to targeted attacks. Given the importance of making high-quality public health data publicly available while preserving the privacy of the underlying data subjects, we aim to improve the utility of a recently developed approach for generating Poisson-distributed, differentially private synthetic data by using publicly available information to truncate the range of the synthetic data. Specifically, we utilize county-level population information from the US Census Bureau and national death reports produced by the CDC to inform prior distributions on county-level death rates and infer reasonable ranges for Poisson-distributed, county-level death counts. In doing so, the requirements for satisfying differential privacy for a given privacy budget can be reduced by several orders of magnitude, thereby leading to substantial improvements in utility. To illustrate our proposed approach, we consider a dataset comprised of over 26,000 cancer-related deaths from the Commonwealth of Pennsylvania belonging to over 47,000 combinations of cause-of-death and demographic variables such as age, race, sex, and county-of-residence and demonstrate the proposed framework’s ability to preserve features such as geographic, urban/rural, and racial disparities present in the true data.