Bayesian disclosure risk assessment: predicting small frequencies in contingency tables

Bayesian disclosure risk assessment: predicting small frequencies in contingency tables
复制标题

DOI:
10.1111/j.1467-9876.2007.00591.x
复制
发表时间:
2007-01-01
影响因子:
1.6
通讯作者:
Webb, Emily L.
Webb, Emily L.
中科院分区:
数学3区
文献类型:
--
作者:
Forster, Jonathan J.;Webb, Emily L.

文献摘要

被引文献

相似文献

我们提出了一种方法来评估个人识别的分类数据的发布中的风险。这需要精确计算列联表中具有小样本频率的那些单元格的预测概率,使得问题与通常的列联表估计有些不同,其中兴趣通常集中在高概率区域。我们的方法是贝叶斯,并提供识别风险的后验预测概率。通过将模型的不确定性纳入我们的分析中,我们可以提供比忽略数据集的多变量结构的方法更现实的单个细胞计数的披露风险估计。
We propose an approach for assessing the risk of individual identification in the release of categorical data. This requires the accurate calculation of predictive probabilities for those cells in a contingency table which have small sample frequencies, making the problem somewhat different from usual contingency table estimation, where interest is generally focused on regions of high probability. Our approach is Bayesian and provides posterior predictive probabilities of identification risk. By incorporating model uncertainty in our analysis, we can provide more realistic estimates of disclosure risk for individual cell counts than are provided by methods which ignore the multivariate structure of the data set.