Accelerating Decision Tree Ensemble with Guided Branch Approximation

Accelerating Decision Tree Ensemble with Guided Branch Approximation
复制标题

使用引导分支逼近加速决策树集成

DOI:
10.1145/3535044.3535048
复制
发表时间:
2022
期刊:
International Symposium on Highly-Efficient Accelerators and Reconfigurable Technologies 2022 (HEART 2022)
影响因子:
--
通讯作者:
Keisuke Kamahori and Shinya Takamaeda-Yamazaki
Keisuke Kamahori and Shinya Takamaeda-Yamazaki
中科院分区:
--
文献类型:
--
作者:
Keisuke Kamahori and Shinya Takamaeda-Yamazaki

文献摘要

相似文献

在低功耗边缘设备上处理轻量级机器学习(ML)算法(如决策树集成(DTE))是有益的;然而,这些设备通常具有有限的资源,并且特定于域的加速器不容易获得。因此,需要针对轻量级嵌入式微控制器上的ML工作负载的节能和资源高效的加速机制,而无需额外的硬件加速器。然而,当在传统的有序流水线处理器上执行DTE时,与分支误预测相关联的惩罚可能是性能瓶颈。本研究提出了引导分支近似(GBA),一种近似计算方法,通过选择性地忽略分支指令的正确性来提高轻量级通用处理器上DTE的性能。GBA通过推测性地执行所选择的分支指令而不对分支误预测进行任何回滚来增强性能。GBA允许程序员和高级ML框架注释近似分支指令,并确保目标应用程序的服务质量(QoS)。GBA包括以下内容:1)近似分支指令格式,一种忽略分支预测器的错误预测的新型分支指令,以及2)基于硬件的QoS机制,其动态地管理可近似分支指令的执行以防止不期望的QoS降级。我们评估所提出的想法,使用软件模拟器的顺序流水线处理器。实验表明,在最佳情况下,只需对硬件进行轻微修改,GBA就可以将总执行时间减少15%以上,同时保留DTE算法的QoS。
Processing lightweight machine learning (ML) algorithms, such as decision tree ensemble (DTE), on low-power edge devices is beneficial; however, these devices usually have limited resources, and domain-specific accelerators are not readily available. Therefore, energy- and resource-efficient acceleration mechanisms for ML workloads on lightweight embedded microcontrollers without additional hardware accelerators are desired. However, the penalties associated with branch mispredictions can be performance bottlenecks when executing DTE on conventional in-order pipelined processors. This study proposes the Guided Branch Approximation (GBA), an approximate computing approach to improve the performance of DTE on lightweight general-purpose processors by selectively ignoring the correctness of branch instructions. GBA enhances the performance by speculatively executing selected branch instructions without any rollback on branch mispredictions. GBA allows programmers and high-level ML frameworks to annotate approximal branch instructions and to ensure target applications’ quality of service (QoS). GBA comprises the following: 1) the approximate branch instruction format, a new type of branch instruction that ignores the wrong prediction of branch predictors, and 2) a hardware-based QoS mechanism that dynamically manages the execution of approximable branch instructions to prevent undesirable QoS degradation. We evaluate the proposed idea on an in-order pipeline processor using a software simulator. Experiments show that GBA can reduce the total execution time by more than 15 % while preserving the QoS of the DTE algorithm in the best-case scenario with a slight modification to the hardware.