Accounting for Intruder Uncertainty Due to Sampling When Estimating Identification Disclosure Risks in Partially Synthetic Data

Accounting for Intruder Uncertainty Due to Sampling When Estimating Identification Disclosure Risks in Partially Synthetic Data
复制标题

在估计部分合成数据中的身份泄露风险时考虑采样引起的入侵者不确定性

DOI:
10.1007/978-3-540-87471-3_19
复制
发表时间:
2008
期刊:
Trans. Data Priv.
影响因子:
--
通讯作者:
Jerome P. Reiter
Jerome P. Reiter
中科院分区:
--
文献类型:
--
作者:
Jörg Drechsler;Jerome P. Reiter

文献摘要

被引文献

相似文献

部分合成数据包括最初调查的单位,其中有些收集的数值,如披露风险高的敏感数值或关键识别特征的数值,被从统计模型中多次提取的数值所取代。由于原始记录保留在文件中,因此入侵者可能能够将这些记录链接到外部数据库,即使值是合成的。我们说明了统计机构如何在发布此类数据之前评估身份披露的风险。当入侵者知道谁在样本中,当入侵者不知道谁在样本中时,我们计算风险度量。我们使用分类和回归树来综合来自美国当前人口调查的数据。
Partially synthetic data comprise the units originally surveyed with some collected values, such as sensitive values at high risk of disclosure or values of key identifiers, replaced with multiple draws from statistical models. Because the original records remain on the file, intruders may be able to link those records to external databases, even though values are synthesized. We illustrate how statistical agencies can evaluate the risks of identification disclosures before releasing such data. We compute risk measures when intruders know who is in the sample and when the intruders do not know who is in the sample. We use classification and regression trees to synthesize data from the U.S. Current Population Survey.