A distributed frequent itemset mining algorithm using Spark for Big Data analytics
A distributed frequent itemset mining algorithm using Spark for Big Data analytics
复制标题
使用 Spark 进行大数据分析的分布式频繁项集挖掘算法
DOI:
10.1007/s10586-015-0477-1
复制
发表时间:
2015-10
影响因子:
4.4
通讯作者:
Ma, Yunlong
中科院分区:
文献类型:
--
作者:
Gui, Feng;Shen, Weiming;Shami, Abdallah;Ma, Yunlong
Frequent itemset mining is an essential step in the process of association rule mining. Conventional approaches for mining frequent itemsets in big data era encounter significant challenges when computing power and memory space are limited. This paper proposes an efficient distributed frequent itemset mining algorithm (DFIMA) which can significantly reduce the amount of candidate itemsets by applying a matrix-based pruning approach. The proposed algorithm has been implemented using Spark to further improve the efficiency of iterative computation. Numeric experiment results using standard benchmark datasets by comparing the proposed algorithm with the existing algorithm, parallel FP-growth, show that DFIMA has better efficiency and scalability. In addition, a case study has been carried out to validate the feasibility of DFIMA.
登录
查看更多内容
DOI:
10.1145/2463676.2465288
发表时间:
2012-11
期刊:
--
影响因子:
--
作者:
Reynold Xin;Josh Rosen;M. Zaharia;M. Franklin;S. Shenker;I. Stoica
通讯作者:
Reynold Xin;Josh Rosen;M. Zaharia;M. Franklin;S. Shenker;I. Stoica
DOI:
10.1007/3-540-36175-8_47
发表时间:
2003-04
期刊:
--
影响因子:
--
作者:
Iko Pramudiono;M. Kitsuregawa
通讯作者:
Iko Pramudiono;M. Kitsuregawa
DOI:
10.1007/s13042-013-0172-6
发表时间:
2013-05
影响因子:
5.6
作者:
M. Mohamed;Mohammed M. Darwieesh
通讯作者:
M. Mohamed;Mohammed M. Darwieesh
DOI:
10.1145/1007730.1007744
发表时间:
2004-06
期刊:
SIGKDD Explor.
影响因子:
--
作者:
Bart Goethals;Mohammed J. Zaki
通讯作者:
Bart Goethals;Mohammed J. Zaki
影响因子:
2.7
作者:
Lamine M. Aouad;Nhien-An Le-Khac;Mohand Tahar Kechadi
通讯作者:
Lamine M. Aouad;Nhien-An Le-Khac;Mohand Tahar Kechadi