A scalable random forest algorithm based on MapReduce

A scalable random forest algorithm based on MapReduce
复制标题

DOI:
10.1109/icsess.2013.6615438
复制
发表时间:
2013-05
期刊:
2013 IEEE 4th International Conference on Software Engineering and Service Science
影响因子:
--
通讯作者:
Jiawei Han;Yanheng Liu;Xin Sun
Jiawei Han;Yanheng Liu;Xin Sun
中科院分区:
其他
文献类型:
--
作者:
Jiawei Han;Yanheng Liu;Xin Sun

文献摘要

被引文献

相似文献

随机森林是一种流行的机器学习数据分类算法。本文提出SMRF算法——一种基于MapReduce模型的改进的可扩展随机森林算法。这种新算法在计算机集群或云计算环境中对海量数据集进行数据分类。 SMRF 通过分布式处理和优化多个参与计算节点上的数据子集。实验结果表明,SMRF算法与传统随机森林算法相比,精度下降同样,但性能更高。 SMRF算法比传统的随机森林算法更适合分布式计算环境中的海量数据集分类。
Random Forest is a popular data classification algorithm for machine learning. This paper proposes SMRF algorithm--an improved scalable Random Forest algorithm based on Map Reduce model. This new algorithm makes data classification in computer cluster or cloud computing environment for massive datasets. SMRF processes and optimizes the subsets of the data across multiple participating computing nodes by distributing. The experimental results show that the SMRF algorithm has the equally accuracy degradation but higher performance while comparing with traditional Random Forest algorithm. SMRF algorithm is more suitable to classify massive data sets in distributing computing environment than traditional Random Forest algorithm.