Hybrid microaggregation for privacy preserving data mining

Hybrid microaggregation for privacy preserving data mining
复制标题

DOI:
10.1007/s12652-018-1122-7
复制
发表时间:
2020-01-01
影响因子:
--
通讯作者:
Perera, Charith
Perera, Charith
中科院分区:
计算机科学3区
文献类型:
--
作者:
Abidi, Balkis;Ben Yahia, Sadok;Perera, Charith

文献摘要

被引文献

相似文献

微聚合的k-匿名是最常用的匿名技术之一。这种成功是由于在信息损失和身份泄露风险之间实现了利益权衡。然而,该方法可能具有一些缺点。在披露限制方面,缺乏对属性披露的保护。在数据实用程序方面,处理真实的数据集是一项具有挑战性的任务。事实上,后者的特点是其大量的属性和噪声数据的存在,这样的离群值,甚至,数据与缺失值。生成对数据挖掘任务有用的匿名个体数据,同时减少噪声数据的影响是一项引人注目的任务。在本文中,我们介绍了一种新的微聚集方法,称为HM-pfsom,基于模糊可能性聚类。我们提出的方法通过混合方式操作。这意味着匿名化过程是按相似数据块应用的。因此,我们可以帮助减少匿名化过程中的信息丢失。HM-pfsom方法提出研究每个子数据集中机密属性的分布。然后,根据后一种分布,以保持匿名化微数据内的机密属性的多样性的方式确定隐私参数k。这允许降低机密信息的泄露风险。
k-Anonymity by microaggregation is one of the most commonly used anonymization techniques. This success is owe to the achievement of a worth of interest trade-off between information loss and identity disclosure risk. However, this method may have some drawbacks. On the disclosure limitation side, there is a lack of protection against attribute disclosure. On the data utility side, dealing with a real datasets is a challenging task to achieve. Indeed, the latter are characterized by their large number of attributes and the presence of noisy data, such that outliers or, even, data with missing values. Generating an anonymous individual data useful for data mining tasks, while decreasing the influence of noisy data is a compelling task to achieve. In this paper, we introduce a new microaggregation method, called HM-pfsom, based on fuzzy possibilistic clustering. Our proposed method operates through an hybrid manner. This means that the anonymization process is applied per block of similar data. Thus, we can help to decrease the information loss during the anonymization process. The HM-pfsom approach proposes to study the distribution of confidential attributes within each sub-dataset. Then, according to the latter distribution, the privacy parameter k is determined, in such a way to preserve the diversity of confidential attributes within the anonymized microdata. This allows to decrease the disclosure risk of confidential information.