Estimating risks of identification disclosure in microdata

Estimating risks of identification disclosure in microdata
复制标题

DOI:
10.1198/016214505000000619
复制
发表时间:
2005-12-01
影响因子:
3.7
通讯作者:
Reiter, JP
Reiter, JP
中科院分区:
数学1区
文献类型:
--
作者:
Reiter, JP

文献摘要

被引文献

相似文献

当统计机构向公众发布微数据时,恶意用户(入侵者)可能会将发布数据中的记录链接到外部数据库中的记录。以无法防止这种身份识别的方式发布数据可能会使机构失去信誉,或者对某些数据构成违法行为。为了限制信息披露,各机构经常发布经过修改的数据版本;然而,通常仍然存在被识别的风险。本文应用并扩展了Duncan和Lambert开发的用于计算采样单元识别概率的框架。它描述了专门针对通过重新编码和顶编码变量、数据交换或添加随机噪声(以及这些常见数据更改技术的组合)更改的数据量身定制的方法,机构可以使用这些方法来评估入侵者的威胁,这些入侵者拥有关于变量之间关系和数据更改方法的信息。本文使用来自当前人口调查的数据,说明了在入侵者知识的不同假设下评估竞争版本的身份披露风险的逐步过程。风险措施提出了个别单位和整个数据集。
When statistical agencies release microdata to the public, malicious users (intruders) may be able to link records in the released data to records in external databases. Releasing data in ways that fail to prevent such identifications may discredit the agency or, for some data, constitute a breach of law. To limit disclosures, agencies often release altered versions of the data; however, there usually remain risks of identification. This article applies and extends the framework developed by Duncan and Lambert for computing probabilities of identification for sampled units. It describes methods tailored specifically to data altered by recoding and topcoding variables, data swapping, or adding random noise (and combinations of these common data alteration techniques) that agencies can use to assess threats from intruders who possess information on relationships among variables and the methods of data alteration. Using data from the Current Population Survey, the article illustrates a step-by-step process for evaluating identification disclosure risks for competing releases under varying assumptions of intruders' knowledge. Risk measures are presented for individual units and for entire datasets.