Protecting respondents' identities in microdata release

Protecting respondents' identities in microdata release
复制标题

DOI:
10.1109/69.971193
复制
发表时间:
2001-11-01
影响因子:
8.9
通讯作者:
Samarati, P
Samarati, P
中科院分区:
计算机科学2区
文献类型:
--
作者:
Samarati, P

文献摘要

被引文献

相似文献

当今的全球网络社会对信息的传播和共享提出了巨大的需求。虽然过去发布的信息大多以表格和统计形式呈现,但现在许多情况需要发布具体数据(微观数据)。为了保护信息所涉及的实体(称为受访者)的匿名性,数据持有者通常会删除或加密显式标识符,例如姓名、地址和电话号码。然而,去识别化数据并不能保证匿名。发布的信息通常包含其他数据,例如种族、出生日期、性别和邮政编码,这些数据可以链接到公开信息,以重新识别受访者并推断不打算披露的信息。在本文中,我们解决了发布微观数据的问题,同时保护数据所涉及的受访者的匿名性。该方法基于 k-匿名性的定义。如果尝试将明确的识别信息链接到其内容,则表提供 k-匿名性,将信息映射到至少 k 个实体。我们说明了如何在不损害通过使用泛化和抑制技术发布的信息的完整性(或真实性)的情况下提供 k-匿名性。我们引入了最小泛化的概念,该概念捕获了发布过程的属性,不会使数据扭曲超过实现 k-匿名所需的程度,并提出了一种用于计算此类泛化的算法。我们还讨论了可能的偏好政策,以在不同的最小概括中进行选择。
Today's globally networked society places great demand on the dissemination and sharing of information. While in the past released information was mostly in tabular and statistical form, many situations call today for the release of specific data (microdata). In order to protect the anonymity of the, entities (called respondents) to which information refers, data holders often remove or encrypt explicit identifiers such as names, addresses, and phone numbers. Deidentifying data, however, provides no guarantee of anonymity. Released information often contains other data, such as race, birth date, sex, and ZIP code, that can be linked to publicly available information to reidentify respondents and inferring information that was not intended for disclosure. In this paper we address the problem of releasing microdata while safeguarding the anonymity of the respondents to which the data refer. The approach is based on the definition of k-anonymity. A table provides k-anonymity if attempts to link explicitly identifying information to its content map the information to at least k entities. We illustrate how k-anonymity can be provided without compromising the Integrity (or truthfulness) of the information released by using generalization and suppression techniques. We introduce the concept of minimal generalization that captures the property of the release process not to distort the data more than needed to achieve k-anonymity, and present an algorithm for the computation of such a generalization. We also discuss possible preference policies to choose among different minimal generalizations.