Semantic diversity: Privacy considering distance between values of sensitive attribute

Semantic diversity: Privacy considering distance between values of sensitive attribute
复制标题

DOI:
10.1016/j.cose.2020.101823
复制
发表时间:
2020-07
期刊:
Comput. Secur.
影响因子:
--
通讯作者:
Keiichiro Oishi;Y. Sei;Yasuyuki Tahara;Akihiko Ohsuga
Keiichiro Oishi;Y. Sei;Yasuyuki Tahara;Akihiko Ohsuga
中科院分区:
其他
文献类型:
--
作者:
Keiichiro Oishi;Y. Sei;Yasuyuki Tahara;Akihiko Ohsuga

文献摘要

被引文献

相似文献

包含个人信息并通过众测收集的数据库可用于各种目的。因此,数据库持有者可能希望与其他组织共享其数据库。然而,由于数据库包含有关个人的信息,数据库接收者必须考虑隐私问题。主流的隐私保护指标之一l-diversity,保证在数据库中识别出个体敏感属性值的概率小于1/l。然而,当在敏感属性中存在若干语义上相似的值时,即使执行匿名化以满足多样性,也有可能不满足实际多样性。例如,如果匿名数据库满足3-多样性,则攻击者可以知道Alice病的候选者是HIV-1(M)、HIV-1(N)和HIV-2的集合。在这种情况下,攻击者可以得出结论,爱丽丝有艾滋病毒,虽然详细的类型仍然未知。在这项研究中,为了解决如何实际的多样性不能考虑现有的l-多样性,我们提出了一种新的隐私指标,(l,d)-语义多样性,和一个算法,匿名的数据库,以满足(l,d)-语义多样性。我们还提出了一个分析算法,是适合所提出的匿名算法,因为匿名算法的输出是难以理解的。我们提出的算法进行了实验评估,使用合成和真实的数据集。
A database that contains personal information and is collected by crowdsensing can be used for various purposes. Therefore, database holders may want to share their databases with other organizations. However, since a database contains information about individuals, database recipients must take privacy concerns into consideration. One of the mainstream privacy protection indicators,l-diversity, guarantees that the probability of identifying a sensitive attribute value of an individual in a database is less than 1/l. However, when there are several semantically similar values in the sensitive attribute, there is a possibility that actual diversity is not satisfied, even if anonymization is performed to satisfyl-diversity. For example, an attacker may know that candidates of Alice’s disease are a set of HIV-1(M), HIV-1(N), and HIV-2 if the anonymized database satisfies 3-diversity. In this case, the attacker can conclude that Alice has HIV, although the detailed type remains unknown. In this research, to solve how actual diversity cannot be taken into consideration with existingl-diversity, we proposed a novel privacy indicator, (l, d)-semantic diversity, and an algorithm that anonymizes a database to satisfy (l, d)-semantic diversity. We also proposed an analysis algorithm that is suitable for the proposed anonymizing algorithm because the output of the anonymizing algorithm is difficult to understand. Our proposed algorithms were experimentally evaluated using synthetic and real datasets.