Accelerating Random Forest Classification on GPU and FPGA

Accelerating Random Forest Classification on GPU and FPGA
复制标题

在 GPU 和 FPGA 上加速随机森林分类

DOI:
10.1145/3545008.3545067
复制
发表时间:
2022
期刊:
ICPP '22: Proceedings of the 51st International Conference on Parallel Processing
影响因子:
--
通讯作者:
Becchi, Michela
Becchi, Michela
中科院分区:
--
文献类型:
--
作者:
Shah, Milan;Neff, Reece;Wu, Hancheng;Minutoli, Marco;Tumeo, Antonino;Becchi, Michela

文献摘要

参考文献

相似文献

随机森林(RF)是一种常用的机器学习方法,用于跨越各种应用领域的分类和回归任务,包括生物信息学,商业分析和软件优化。虽然以前的工作主要集中在提高RF的训练性能,但许多应用,如恶意软件识别,癌症预测和银行欺诈检测,都需要快速RF分类。在这项工作中,我们在GPU和FPGA上加速RF分类。为了提供对大数据集的有效支持,我们提出了一种适合GPU/FPGA存储器层次结构的分层存储器布局。我们设计了三个RF分类代码变体的基础上,布局,我们调查这些内核的GPU和FPGA特定的考虑。我们在Nvidia Xp GPU和Xilinx Alveo U250 FPGA加速器卡上进行的实验评估使用了数百万样本和数十个功能的公开数据集,涵盖了各个方面。首先,我们评估我们的分层数据结构的性能优势,在标准的压缩稀疏行(CSR)格式。其次,我们将我们的GPU实现与cuML进行比较,cuML是一个针对Nvidia GPU的机器学习库。第三,我们探讨了在RF中使用不同树深度所导致的性能/精度权衡。最后,我们对GPU和FPGA实现进行了比较性能分析。我们的评估表明,在GPU上报告最佳性能的同时,我们的代码变体在GPU和FPGA上的性能都优于CSR基线。对于高精度目标,我们的GPU实现比CSR产生5-9倍的加速,比Nvidia的cuML库产生高达2倍的加速。
Random Forests (RFs) are a commonly used machine learning method for classification and regression tasks spanning a variety of application domains, including bioinformatics, business analytics, and software optimization. While prior work has focused primarily on improving performance of the training of RFs, many applications, such as malware identification, cancer prediction, and banking fraud detection, require fast RF classification.In this work, we accelerate RF classification on GPU and FPGA. In order to provide efficient support for large datasets, we propose a hierarchical memory layout suitable to the GPU/FPGA memory hierarchy. We design three RF classification code variants based on that layout, and we investigate GPU- and FPGA-specific considerations for these kernels. Our experimental evaluation, performed on an Nvidia Xp GPU and on a Xilinx Alveo U250 FPGA accelerator card using publicly available datasets on the scale of millions of samples and tens of features, covers various aspects. First, we evaluate the performance benefits of our hierarchical data structure over the standard compressed sparse row (CSR) format. Second, we compare our GPU implementation with cuML, a machine learning library targeting Nvidia GPUs. Third, we explore the performance/accuracy tradeoff resulting from the use of different tree depths in the RF. Finally, we perform a comparative performance analysis of our GPU and FPGA implementations. Our evaluation shows that, while reporting the best performance on GPU, our code variants outperform the CSR baseline both on GPU and FPGA. For high accuracy targets, our GPU implementation yields a 5-9 × speedup over CSR, and up to a 2 × speedup over Nvidia’s cuML library.
使用 Altera SDK for OpenCL 加速随机森林分类
DOI: --
发表时间: 2016
期刊: International Conference on Field-Programmable Technology
影响因子: --
作者:
Hiroki Nakahara;Akira Jinguji;Tomonori Fujii;S. Sato
通讯作者: S. Sato
DOI: --
发表时间: 2013
期刊: 2013 SC - International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子: --
作者:
Michael Goldfarb;Youngjoon Jo;Milind Kulkarni
通讯作者: Milind Kulkarni
一种新型的基于 FPGA 的二叉搜索树高吞吐量加速器
DOI: --
发表时间: 2019
期刊: International Symposium on High Performance Computing Systems and Applications
影响因子: --
作者:
Oyku Melikoglu;Oğuz Ergin;Behzad Salami;Julián Pavón;O. Unsal;A. Cristal
通讯作者: A. Cristal
DOI: --
发表时间: 2017
期刊: International Conference on Parallel and Distributed Systems
影响因子: --
作者:
Hancheng Wu;M. Becchi
通讯作者: M. Becchi