Publishing Time-Series Data under Preservation of Privacy and Distance Orders

Publishing Time-Series Data under Preservation of Privacy and Distance Orders
复制标题

在保护隐私和距离顺序的情况下发布时间序列数据

DOI:
--
复制
发表时间:
2010
期刊:
International Conference on Database and Expert Systems Applications
影响因子:
--
通讯作者:
E. Bertino
E. Bertino
中科院分区:
--
文献类型:
--
作者:
Yang;Hea;Sang;E. Bertino

文献摘要

被引文献

相似文献

在本文中,我们解决了在发布敏感的时间序列数据时保持挖掘精度和隐私的问题。例如,患有心脏病的人不想公开他们的心电图时间序列,但他们仍然允许从他们的时间序列中挖掘一些准确的模式。在此基础上,我们引入了相关的假设和要求,证明了只有随机化方法满足所有的假设,但即使是这些方法也不满足要求。因此,我们讨论了满足所有假设和要求的基于随机化的解决方案。为此,我们使用分段聚集近似(PAA)的噪声平均效应,这可能会减轻破坏随机扰动时间序列的距离顺序的问题。基于噪声平均效应,我们首先提出了两个简单的解决方案,使用随机数据扰动发布时间序列,同时利用PAA距离计算距离。然而,这两种解决方案之间的不确定性和距离顺序的权衡。因此,我们提出了两个更先进的解决方案,利用这两个天真的解决方案。实验结果表明,我们的先进的解决方案是上级优于朴素的解决方案。
In this paper we address the problem of preserving mining accuracy as well as privacy in publishing sensitive time-series data. For example, people with heart disease do not want to disclose their electrocardiogram time-series, but they still allow mining of some accurate patterns from their time-series. Based on this observation, we introduce the related assumptions and requirements.We show that only randomization methods satisfy all assumptions, but even those methods do not satisfy the requirements. Thus, we discuss the randomization-based solutions that satisfy all assumptions and requirements. For this purpose, we use the noise averaging effect of piecewise aggregate approximation (PAA), which may alleviate the problem of destroying distance orders in randomly perturbed time-series. Based on the noise averaging effect, we first propose two naive solutions that use the random data perturbation in publishing time-series while exploiting the PAA distance in computing distances. There is, however, a tradeoff between these two solutions with respect to uncertainty and distance orders. We thus propose two more advanced solutions that take advantages of both naive solutions. Experimental results show that our advanced solutions are superior to the naive solutions.