A Probabilistic Procedure for Anonymisation and Analysis of Perturbed Datasets
A Probabilistic Procedure for Anonymisation and Analysis of Perturbed Datasets
复制标题
扰动数据集匿名化和分析的概率过程
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
H. Goldstein
中科院分区:
文献类型:
--
作者:
H. Goldstein
The requirement to anonymise datasets that are to be released for secondary analysis needs to be balanced by the need to allow their analysis to provide efficient and consistent parameter estimates. The proposal in the present paper is to use the addition of random noise to some or all variables in a released (already pseudonymised) data set where the values of some identifying variables for individuals of interest are also available to an external ‘attacker’ who wishes to identify those individuals so that they can interrogate their records in the dataset. To avoid such identification enough noise needs to be generated and added to these identifying variables. The noise so generated then needs to be accounted for at the analysis stage to provide required parameter estimates. Where the characteristics of the noise are made available to the analyst by the agency providing the data, we propose a method that allows a valid analysis. This is formally a measurement error model and there exist procedures for model fitting that recovers consistent estimates of the true model parameters. The paper shows how an appropriate noise distribution can be determined and at the analysis stage describes a Bayesian MCMC algorithm that allows for noise removal.
DOI:
10.1214/09-aoas317
发表时间:
2010
期刊:
The Annals of Applied Statistics
影响因子:
--
作者:
Shlomo N
通讯作者:
Shlomo N