Anonymization of Sensitive Quasi-Identifiers for l-Diversity and t-Closeness

Anonymization of Sensitive Quasi-Identifiers for l-Diversity and t-Closeness
复制标题

DOI:
10.1109/tdsc.2017.2698472
复制
发表时间:
2019-07-01
影响因子:
7.3
通讯作者:
Ohsuga, Akihiko
Ohsuga, Akihiko
中科院分区:
计算机科学2区
文献类型:
--
作者:
Sei, Yuichi;Okumura, Hiroshi;Ohsuga, Akihiko

文献摘要

被引文献

相似文献

许多研究隐私保护数据挖掘已经提出。他们中的大多数人认为,他们可以分离准标识符(QID)的敏感属性。例如,他们假设地址、工作和年龄是QID,但不是敏感属性,疾病名称是敏感属性,但不是QID。然而,所有这些属性在实践中可以具有既是敏感属性又是QID的特征。在本文中,我们把这些属性作为敏感的QID,我们提出了新的隐私模型,即(l1,.,lq)-分集和(t1,.,tq)-接近度,以及可以处理敏感QID的方法。我们的方法是由两个算法:匿名化算法和重建算法。由数据持有者进行的匿名化算法简单但有效,而由数据分析者进行的重建算法可以根据每个数据分析者的目标进行。我们提出的方法进行了实验评估,使用真实的数据集。
A number of studies on privacy-preserving data mining have been proposed. Most of them assume that they can separate quasi-identifiers (QIDs) from sensitive attributes. For instance, they assume that address, job, and age are QIDs but are not sensitive attributes and that a disease name is a sensitive attribute but is not a QID. However, all of these attributes can have features that are both sensitive attributes and QIDs in practice. In this paper, we refer to these attributes as sensitive QIDs and we propose novel privacy models, namely, (l1, ... , lq)-diversity and (t1, ... , tq)-closeness, and a method that can treat sensitive QIDs. Our method is composed of two algorithms: An anonymization algorithm and a reconstruction algorithm. The anonymization algorithm, which is conducted by data holders, is simple but effective, whereas the reconstruction algorithm, which is conducted by data analyzers, can be conducted according to each data analyzer's objective. Our proposed method was experimentally evaluated using real data sets.