Estimation of cost of k-anonymity in the number of dummy records

Estimation of cost of k-anonymity in the number of dummy records
复制标题

虚拟记录数量中 k 匿名成本的估计

DOI:
10.1007/s12652-021-03369-5
复制
发表时间:
2022
影响因子:
--
通讯作者:
Kikuchi Hiroaki
Kikuchi Hiroaki
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ito Satoshi;Kikuchi Hiroaki

文献摘要

参考文献

相似文献

去身份化是通过处理个人身份信息,防止个人从原始交易数据中被识别出来的过程。k-匿名化是去身份化的代表方法之一,它处理数据,使至少k个用户具有相同的记录。匿名化的方法之一是在数据中添加虚拟记录,以保护具有独特历史的用户。该方法的成本分叉匿名是原始数据和处理后数据的记录数之差,只有在确定参数k和处理数据后才能计算出成本分叉匿名。然而,我们希望在处理之前计算成本并找到k的最佳值,因为处理各种各样的大数据是非常昂贵的。在本文中,我们提出了一个新的模型的交易数据,给我们一个概率分布和期望值的值在数据的假设下,所有的值独立发生的概率一致。应用我们的数据模型,即使在处理之前,也可以评估k-匿名数据的成本。
De-identification is a process to prevent individuals from being identified from original transaction data by processing personal identification information.k-anonymization, which processes data so that at leastkusers have the same records, is one of the representative methods of de-identification. One of the methods ofk-anonymization is adding dummy records into the data to protect users who have unique histories. For this method, the cost fork-anonymization is the difference in the number of records between the original data and the processed data, and it can be calculated only after deciding the parameterkand processing data. However, we want to calculate the cost before processing and find the optimal value ofkbecause processing the big data with variouskis very costly. In this paper, we propose a new model of transaction data that gives us a probability distribution and an expected value of values in data under the assumption that all values occur independently with uniform probability. Applying our data model, it is possible to evaluate the cost ofk-anonymized data even before processing.
轨迹数据隐私风险模型
DOI: 10.1186/s40493-015-0020-6
发表时间: 2015
期刊: Journal of Trust Management
影响因子: --
作者:
A. Basu;A. Monreale;R. Trasarti;J. Corena;F. Giannotti;D. Pedreschi;S. Kiyomoto;Yutaka Miyake;Tadashi Yanagihara
通讯作者: Tadashi Yanagihara