Cludoop: An Efficient Distributed Density-Based Clustering for Big Data Using Hadoop

Cludoop: An Efficient Distributed Density-Based Clustering for Big Data Using Hadoop
复制标题

DOI:
10.1155/2015/579391
复制
发表时间:
2015-06
影响因子:
2.3
通讯作者:
Yanwei Yu;Jindong Zhao;Xiaodong Wang;Qin Wang;Yonggang Zhang
Yanwei Yu;Jindong Zhao;Xiaodong Wang;Qin Wang;Yonggang Zhang
中科院分区:
计算机科学4区
文献类型:
--
作者:
Yanwei Yu;Jindong Zhao;Xiaodong Wang;Qin Wang;Yonggang Zhang

文献摘要

相似文献

基于密度的大数据集群对于从互联网数据处理到大规模移动对象管理的许多现代应用程序至关重要。本文提出了 Cludooop 算法,这是一种使用 Hadoop 的基于分布式密度的高效大数据聚类方法。首先,我们提出了一种串行聚类算法 CluC,利用单元划分优化和 c-cluster 来快速查找聚类。 CluC利用点周围连通单元的关系来完成点的分类,而不是昂贵的完成邻居查询,这大大减少了距离计算的次数。其次,我们提出了 Cludooop,它可以使用 Map/Reduce 平台上现有的数据分区高效地并行集群超大规模数据。它采用所提出的串行聚类 CluC 作为并行映射器上的插入式聚类,以及单元描述而不是传输中的完整单元,以减少网络和 I/O 成本。在提出的基于单元的原则的指导下,我们还设计了一个合并-细化-合并三步框架,将 c 簇合并到减速器上分配的预聚类结果的叠加上。最后,我们使用大量真实数据和合成数据对 10 台联网商用 PC 进行了全面的实验评估,证明了(1)我们的算法在找到任意形状的正确簇方面的有效性;(2)我们提出的算法比最先进的方法表现出更好的可扩展性和效率。
Density-based clustering for big data is critical for many modern applications ranging from Internet data processing to massive-scale moving object management. This paper proposes Cludoop algorithm, an efficient distributed density-based clustering for big data using Hadoop. First, we propose a serial clustering algorithm CluC by leveraging cell partition optimization and c-cluster to fast find clusters. CluC completes classification of the points using the relationships of connected cells around points instead of expensive completed neighbor query, which significantly reduce the number of distance calculations. Second, we propose the Cludoop, which can efficiently cluster very-large-scale data in parallel using already existing data partition on Map/Reduce platform. It employs the proposed serial clustering CluC as a plugged-in clustering on parallel mapper, along with a cell description instead of completed cell in transmission to reduce both network and I/O costs. Guided by proposed cell-based principles, we also design a Merging-Refinement-Merging 3-step framework to merge c-clusters on the overlay of assigned preclustering result on reducer. Finally, our comprehensive experimental evaluation on 10 network-connected commercial PCs, using both huge-volume real and synthetic data, demonstrates (1) the effectiveness of our algorithm in finding correct clusters with arbitrary shape and (2) the fact that our proposed algorithm exhibits better scalability and efficiency than state-of-the-art method.