An acceleration of a random forest classification using Altera SDK for OpenCL

An acceleration of a random forest classification using Altera SDK for OpenCL
复制标题

使用 Altera SDK for OpenCL 加速随机森林分类

DOI:
--
复制
发表时间:
2016
期刊:
International Conference on Field-Programmable Technology
影响因子:
--
通讯作者:
S. Sato
S. Sato
中科院分区:
--
文献类型:
--
作者:
Hiroki Nakahara;Akira Jinguji;Tomonori Fujii;S. Sato

文献摘要

被引文献

相似文献

随机森林(RF)是一种用于分类和回归的集成机器学习算法。它由随机抽样数据构建的多个决策树组成。与其他机器学习算法相比,RF具有简单、快速的学习和识别能力。它被广泛应用于各种识别系统。由于需要对每棵树进行不平衡跟踪,并且需要对所有树进行通信,因此随机森林不适合gpu等SIMD架构。虽然已经提出了使用FPGA的加速器,但这些实现都是基于HDL设计的。因此,与基于软件的实现相比,它们需要更长的设计时间。在本文中,我们展示了使用Altera SDK for OpenCL的射频加速器,这是一种高级合成。为了加速射频分类,我们提出全流水线架构,利用FPGA上的片上存储器来增加内存带宽。此外,我们采用适当的位定点表示法,而不是32位浮点表示法,以减少硬件尺寸,功耗,并增加内存带宽。我们在Terasic公司的DE5-NET FPGA板上实现了射频,并与CPU和GPU的实现进行了比较,在LPS(每秒查找次数)方面,FPGA的实现比GPU快10.7倍,比CPU快14.0倍。在单位功耗LPS方面,FPGA实现比GPU高61.3倍,比CPU高12.1倍。
A random forest (RF) is a kind of ensemble machine learning algorithm used for a classification and a regression. It consists of multiple decision trees that are built from randomly sampled data. The RF has a simple, fast learning, and identification capability compared with other machine learning algorithms. It is widely used for applicable to various recognition systems. Since it is necessary to un-balanced trace for each tree and requires communication for all the ones, the random forest is not suitable in SIMD architectures such as GPUs. Although the accelerators using the FPGA have been proposed, such implementations were based on HDL design. Thus, they required longer design time compared with the soft-ware based realizations. In this paper, we show the accelerator for the RF using the Altera SDK for OpenCL, which is a kind of high-level synthesis. To accelerate the RF classification, we propose the fully pipelined architecture to increase the memory bandwidth using on-chip memories on the FPGA. Also, we apply appropriate bit fixed point representation instead of 32 bit floating point one in order to reduce the hardware size, power consumption, and increase the memory bandwidth. We implemented the RF on the Terasic Corp. DE5-NET FPGA board, and compared with the CPU and the GPU implementations, As for the LPS (lookups per second), the FPGA realization was 10.7 times faster than the GPU one, and it was 14.0 times faster than the CPU one. As for the LPS per power consumption, the FPGA realization was 61.3 times better than the GPU one, and it was 12.1 times better than the CPU one.