Accounting for Intruder Uncertainty Due to Sampling When Estimating Identification Disclosure Risks in Partially Synthetic Data
Accounting for Intruder Uncertainty Due to Sampling When Estimating Identification Disclosure Risks in Partially Synthetic Data
复制标题
在估计部分合成数据中的身份泄露风险时考虑采样引起的入侵者不确定性
DOI:
10.1007/978-3-540-87471-3_19
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
Jerome P. Reiter
中科院分区:
文献类型:
--
作者:
Jörg Drechsler;Jerome P. Reiter
Partially synthetic data comprise the units originally surveyed with some collected values, such as sensitive values at high risk of disclosure or values of key identifiers, replaced with multiple draws from statistical models. Because the original records remain on the file, intruders may be able to link those records to external databases, even though values are synthesized. We illustrate how statistical agencies can evaluate the risks of identification disclosures before releasing such data. We compute risk measures when intruders know who is in the sample and when the intruders do not know who is in the sample. We use classification and regression trees to synthesize data from the U.S. Current Population Survey.