Preserving confidentiality of high-dimensional tabulated data: Statistical and computational issues

Preserving confidentiality of high-dimensional tabulated data: Statistical and computational issues
复制标题

保护高维表格数据的机密性:统计和计算问题

DOI:
10.1023/a:1025671023941
复制
发表时间:
2003
影响因子:
2.2
通讯作者:
A. Sanil
A. Sanil
中科院分区:
数学2区
文献类型:
--
作者:
A. Dobra;A. Karr;A. Sanil

文献摘要

被引文献

相似文献

传播由机密数据构成的大型列联表所衍生的信息是统计机构的一项主要职责。在本文中,我们针对从单个基础表传播交叉列表(边际子表)时出现的几个计算和算法问题提出了解决方案。这些包括利用稀疏性来支持边际高效计算的数据结构,以及诸如迭代比例拟合之类的算法,还有一种广义形式的穿梭算法,该算法可根据任意一组已发布的边际计算完整表格中(小的、对保密性有威胁的)单元格的精确界限。我们给出了说明这些技术的示例。
Dissemination of information derived from large contingency tables formed from confidential data is a major responsibility of statistical agencies. In this paper we present solutions to several computational and algorithmic problems that arise in the dissemination of cross-tabulations (marginal sub-tables) from a single underlying table. These include data structures that exploit sparsity to support efficient computation of marginals and algorithms such as iterative proportional fitting, as well as a generalized form of the shuttle algorithm that computes sharp bounds on (small, confidentiality threatening) cells in the full table from arbitrary sets of released marginals. We give examples illustrating the techniques.