A Probabilistic Procedure for Anonymisation and Analysis of Perturbed Datasets

A Probabilistic Procedure for Anonymisation and Analysis of Perturbed Datasets
复制标题

扰动数据集匿名化和分析的概率过程

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
H. Goldstein
H. Goldstein
中科院分区:
--
文献类型:
--
作者:
H. Goldstein

文献摘要

参考文献

被引文献

相似文献

对拟发布用于二次分析的匿名数据集的要求需要与允许其分析提供有效和一致的参数估计的需要相平衡。本论文的建议是使用随机噪声添加到发布的(已经被加密的)数据集中的一些或所有变量中,其中感兴趣的个体的一些识别变量的值也可用于希望识别这些个体的外部“攻击者”,以便他们可以询问数据集中的记录。为了避免这种识别,需要生成足够的噪声并将其添加到这些识别变量中。然后需要在分析阶段考虑这样产生的噪声,以提供所需的参数估计。当提供数据的机构向分析师提供噪声特性时,我们提出了一种允许进行有效分析的方法。这在形式上是一个测量误差模型,存在模型拟合的程序,可以恢复真实模型参数的一致估计值。本文展示了如何确定适当的噪声分布,并在分析阶段描述了贝叶斯MCMC算法,允许噪声去除。
The requirement to anonymise datasets that are to be released for secondary analysis needs to be balanced by the need to allow their analysis to provide efficient and consistent parameter estimates. The proposal in the present paper is to use the addition of random noise to some or all variables in a released (already pseudonymised) data set where the values of some identifying variables for individuals of interest are also available to an external ‘attacker’ who wishes to identify those individuals so that they can interrogate their records in the dataset. To avoid such identification enough noise needs to be generated and added to these identifying variables. The noise so generated then needs to be accounted for at the analysis stage to provide required parameter estimates. Where the characteristics of the noise are made available to the analyst by the agency providing the data, we propose a method that allows a valid analysis. This is formally a measurement error model and there exist procedures for model fitting that recovers consistent estimates of the true model parameters. The paper shows how an appropriate noise distribution can be determined and at the analysis stage describes a Bayesian MCMC algorithm that allows for noise removal.
评估基于错误分类的调查微观数据披露限制方法所提供的保护
DOI: 10.1214/09-aoas317
发表时间: 2010
期刊: The Annals of Applied Statistics
影响因子: --
作者:
Shlomo N
通讯作者: Shlomo N