Accelerating Random Forest Classification on GPU and FPGA
Accelerating Random Forest Classification on GPU and FPGA
复制标题
在 GPU 和 FPGA 上加速随机森林分类
DOI:
10.1145/3545008.3545067
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Becchi, Michela
中科院分区:
文献类型:
--
作者:
Shah, Milan;Neff, Reece;Wu, Hancheng;Minutoli, Marco;Tumeo, Antonino;Becchi, Michela
Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification.In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that, while reporting the best performance on GPU, our code variants outperform the CSR baseline both on GPU and FPGA. For high accuracy targets, our GPU implementation yields a 5-9 × speedup over CSR, and up to a 2 × speedup over Nvidia’s cuML library.
登录
查看更多内容
DOI:
--
发表时间:
2016
期刊:
International Conference on Field-Programmable Technology
影响因子:
--
作者:
Hiroki Nakahara;Akira Jinguji;Tomonori Fujii;S. Sato
通讯作者:
S. Sato
DOI:
--
发表时间:
2013
期刊:
2013 SC - International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子:
--
作者:
Michael Goldfarb;Youngjoon Jo;Milind Kulkarni
通讯作者:
Milind Kulkarni
DOI:
--
发表时间:
2019
期刊:
International Symposium on High Performance Computing Systems and Applications
影响因子:
--
作者:
Oyku Melikoglu;Oğuz Ergin;Behzad Salami;Julián Pavón;O. Unsal;A. Cristal
通讯作者:
A. Cristal
DOI:
--
发表时间:
2017
期刊:
International Conference on Parallel and Distributed Systems
影响因子:
--
作者:
Hancheng Wu;M. Becchi
通讯作者:
M. Becchi