A Parallel Multiobjective PSO Weighted Average Clustering Algorithm Based on Apache Spark.

A Parallel Multiobjective PSO Weighted Average Clustering Algorithm Based on Apache Spark.
复制标题

DOI:
10.3390/e25020259
复制
发表时间:
2023-01-31
期刊:
影响因子:
2.7
通讯作者:
Liu, Zhenyu
Liu, Zhenyu
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Ling, Huidong;Zhu, Xinmu;Zhu, Tao;Nie, Mingxing;Liu, Zhenghai;Liu, Zhenyu

文献摘要

参考文献

被引文献

相似文献

基于粒子群优化的多目标聚类算法已经在一些应用中得到了成功的应用。然而,现有的算法都是在单机上实现的,不能直接在集群上并行化,这使得现有算法难以处理大规模数据。随着分布式并行计算框架的发展,数据并行被提出。但是并行度的提高会导致数据分布不均衡的问题,影响聚类效果。提出了一种基于Apache Spark的并行多目标PSO加权平均聚类算法(Spark-MOPSO-Avg)。首先,整个数据集被划分为多个分区,并使用Apache Spark的分布式并行和基于内存的计算缓存在内存中。根据分区中的数据并行计算粒子的局部适应度值。计算完成后,只传输粒子信息,各节点之间不需要传输大量的数据对象,减少了网络中数据的通信,从而有效减少了算法的运行时间。其次,对局部适应度值进行加权平均计算,以改善数据分布不均衡影响结果的问题。实验结果表明,Spark-MOPSO-Avg算法在数据并行下实现了较低的信息损失,损失约1%~ 9%的准确率,但能有效降低算法的时间开销。在Spark分布式集群环境下显示出良好的执行效率和并行计算能力。
Multiobjective clustering algorithm using particle swarm optimization has been applied successfully in some applications. However, existing algorithms are implemented on a single machine and cannot be directly parallelized on a cluster, which makes it difficult for existing algorithms to handle large-scale data. With the development of distributed parallel computing framework, data parallelism was proposed. However, the increase in parallelism will lead to the problem of unbalanced data distribution affecting the clustering effect. In this paper, we propose a parallel multiobjective PSO weighted average clustering algorithm based on apache Spark (Spark-MOPSO-Avg). First, the entire data set is divided into multiple partitions and cached in memory using the distributed parallel and memory-based computing of Apache Spark. The local fitness value of the particle is calculated in parallel according to the data in the partition. After the calculation is completed, only particle information is transmitted, and there is no need to transmit a large number of data objects between each node, reducing the communication of data in the network and thus effectively reducing the algorithm’s running time. Second, a weighted average calculation of the local fitness values is performed to improve the problem of unbalanced data distribution affecting the results. Experimental results show that the Spark-MOPSO-Avg algorithm achieves lower information loss under data parallelism, losing about 1% to 9% accuracy, but can effectively reduce the algorithm time overhead. It shows good execution efficiency and parallel computing capability under the Spark distributed cluster.
DOI: 10.1371/journal.pone.0130995
发表时间: 2015
期刊: PloS one
影响因子: 3.7
作者:
Abubaker A;Baharum A;Alrefaei M
通讯作者: Alrefaei M
DOI: 10.1145/2742642
发表时间: 2015-07-01
影响因子: 16.6
作者:
Mukhopadhyay, Anirban;Maulik, Ujjwal;Bandyopadhyay, Sanghamitra
通讯作者: Bandyopadhyay, Sanghamitra
DOI: 10.1109/tevc.2004.826067
发表时间: 2004-06-01
影响因子: 14.3
作者:
Coello, CAC;Pulido, GT;Lechuga, MS
通讯作者: Lechuga, MS
DOI: 10.1371/journal.pone.0188815
发表时间: 2017
期刊: PloS one
影响因子: 3.7
作者:
Gong C;Chen H;He W;Zhang Z
通讯作者: Zhang Z
在不同的非生物应力下,选择合适的参考基因用于Salix Matsudana中定量实时PCR基因表达分析。
DOI: 10.1038/srep40290
发表时间: 2017-01-25
期刊: Scientific reports
影响因子: 4.6
作者:
Zhang Y;Han X;Chen S;Zheng L;He X;Liu M;Qiao G;Wang Y;Zhuo R
通讯作者: Zhuo R