Clustering Data in Secured, Distributed Datasets

Clustering Data in Secured, Distributed Datasets
复制标题

DOI:
10.1007/978-3-030-24311-1_40
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Sayantan Dey;Lee Carraher;Anindya Moitra;P. Wilsey
Sayantan Dey;Lee Carraher;Anindya Moitra;P. Wilsey
中科院分区:
其他
文献类型:
--
作者:
Sayantan Dey;Lee Carraher;Anindya Moitra;P. Wilsey

文献摘要

相似文献

数据生成和收集的大量增长使得开发机械化方法来分析和提取信息的必要性变得尤为重要。数据聚类是从数据中发现新见解的基本模式之一。然而,高维数据也有其自身的挑战,许多传统的聚类算法在准确性或可扩展性方面都失败了。使问题进一步复杂化的是,敏感数据的不同子集可能驻留在地理位置不同的位置,而数据的敏感性质阻止(或抑制)其访问以进行机械化分析。因此,必须找到从这些安全的分布式数据集的整体中发现信息并保持数据完整性的方法。在本文中,我们开发和评估了一种分布式算法,该算法可以对地理上分离的数据进行聚类,同时保留不共享受保护的高维数据的严格隐私要求。我们在基于分布式映射缩减的平台 Spark 上实现我们的算法,并通过将其与标准数据聚类算法进行比较来展示其性能。
The massive growth in data generation and collection has brought to the forefront the necessity to develop mechanized methods to analyze and extract information from them. Data clustering is one of the fundamental modes to discover new insights from data. However, high dimensional data has its own challenges where many conventional clustering algorithms fails either in accuracy or scalability. To further complicate the issue, distinct subsets of sensitive data may reside in geographically separated locations with the sensitive nature of the data preventing (or inhibiting) its access for mechanized analysis. Thus, methods to discover information from the collective whole of these secured, distributed data sets that also preserves the integrity of the data must be found. In this paper we develop and assess a distributed algorithm that can cluster geographically separated data while simultaneously preserving the strict privacy requirements of non sharing of protected high dimensional data. We implement our algorithm on the distributed map-reduce based platform Spark and demonstrate its performance by comparing it to the standard data clustering algorithms.