Estimating Identification Disclosure Risk Using Mixed Membership Models.

Estimating Identification Disclosure Risk Using Mixed Membership Models.
复制标题

DOI:
10.1080/01621459.2012.710508
复制
发表时间:
2012-12-01
影响因子:
3.7
通讯作者:
Reiter JP
Reiter JP
中科院分区:
数学1区
文献类型:
--
作者:
Manrique-Vallier D;Reiter JP

文献摘要

参考文献

被引文献

相似文献

传播数据的统计机构和其他组织有义务保护数据主体的机密性。例如,恶意的个人可能通过匹配共同特征(关键字)将数据主体与其他数据库中的记录联系起来。成功的链接对于具有在总体中唯一的键组合的数据主体来说特别成问题。因此,作为披露风险评估的一部分,许多数据管理员估计离散键集合上的样本唯一性也是这些键上的总体唯一性的概率。这通常是使用键上的对数线性建模来完成的。然而,对数线性模型可以产生有偏估计的单元格概率稀疏列联表与许多零计数,这往往发生在数据库中有许多关键字。这种偏见可能导致对独特性概率的不可靠估计,从而导致对披露风险的误报。我们提出了一种替代对数线性模型的数据集稀疏键的基础上贝叶斯版本的等级的成员(GoM)模型。我们提出了一个贝叶斯GoM模型的多项变量,并提供了一个MCMC算法来拟合模型。我们评估的方法,从最近的美国人口普查局的公共使用微观数据样本作为一个人口的数据处理,从人口简单的随机样本,和基准估计概率的独特性对人口的价值。与对数线性模型相比,GoM模型提供了对样本中唯一性总数的更准确估计。此外,它们提供了记录级的独特性预测,这些预测主导了基于对数线性模型的预测。
Statistical agencies and other organizations that disseminate data are obligated to protect data subjects’ confidentiality. For example, ill-intentioned individuals might link data subjects to records in other databases by matching on common characteristics (keys). Successful links are particularly problematic for data subjects with combinations of keys that are unique in the population. Hence, as part of their assessments of disclosure risks, many data stewards estimate the probabilities that sample uniques on sets of discrete keys are also population uniques on those keys. This is typically done using log-linear modeling on the keys. However, log-linear models can yield biased estimates of cell probabilities for sparse contingency tables with many zero counts, which often occurs in databases with many keys. This bias can result in unreliable estimates of probabilities of uniqueness and, hence, misrepresentations of disclosure risks. We propose an alternative to log-linear models for datasets with sparse keys based on a Bayesian version of grade of membership (GoM) models. We present a Bayesian GoM model for multinomial variables and offer an MCMC algorithm for fitting the model. We evaluate the approach by treating data from a recent US Census Bureau public use microdata sample as a population, taking simple random samples from that population, and benchmarking estimated probabilities of uniqueness against population values. Compared to log-linear models, GoM models provide more accurate estimates of the total number of uniques in the samples. Additionally, they offer record-level predictions of uniqueness that dominate those based on log-linear models.
DOI: 10.1073/pnas.0307760101
发表时间: 2004-04-06
影响因子: 11.1
作者:
Erosheva, E;Fienberg, S;Lafferty, J
通讯作者: Lafferty, J
DOI: 10.1016/j.jsc.2005.04.003
发表时间: 2006-02-01
影响因子: 0.7
作者:
Eriksson, N;Fienberg, SE;Sullivant, S
通讯作者: Sullivant, S
DOI: 10.1111/j.1467-9876.2007.00591.x
发表时间: 2007-01-01
影响因子: 1.6
作者:
Forster, Jonathan J.;Webb, Emily L.
通讯作者: Webb, Emily L.
DOI: 10.2307/2334349
发表时间: 1974-01-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
GOODMAN, LA
通讯作者: GOODMAN, LA
DOI: 10.1214/07-aoas126
发表时间: 2007-12-01
影响因子: 1.8
作者:
Erosheva, Elena A.;Fienberg, Stephen E.;Joutard, Cyrille
通讯作者: Joutard, Cyrille