HERALD: Optimizing Heterogeneous DNN Accelerators for Edge Devices

HERALD: Optimizing Heterogeneous DNN Accelerators for Edge Devices
复制标题

DOI:
--
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Hyoukjun Kwon;Liangzhen Lai;T. Krishna;V. Chandra
Hyoukjun Kwon;Liangzhen Lai;T. Krishna;V. Chandra
中科院分区:
其他
文献类型:
--
作者:
Hyoukjun Kwon;Liangzhen Lai;T. Krishna;V. Chandra

文献摘要

被引文献

相似文献

诸如基于深度神经网络(DNN)且具有多个子任务(例如图像分割、手部追踪等)的虚拟现实(VR)等新型实时应用正在兴起。此类应用的每个子任务都依赖多个不同的DNN,并且需要满足每个子任务的目标处理速率。因此,新的复合DNN工作负载在两个方面对加速器设计提出了新的挑战:(1)满足子任务中每个DNN的目标处理速率;(2)高效处理异构层。作为一种解决方案,我们探索异构DNN加速器(HDA)。HDA由多个加速器基板(即子加速器)组成,以支持模型并行性,其实现不同的映射方式,从而为异构DNN层提供适应性。然而,HDA的性能和能耗在很大程度上取决于(1)我们如何为子加速器划分硬件资源,以及(2)我们如何在子加速器上调度层。因此,我们提出了一个HDA优化框架——Herald,它对硬件划分和层调度进行协同优化。Herald利用DNN层依赖图的简单性来降低问题的复杂性,在配备i9 - 9880H处理器和16GB内存的笔记本电脑上,平均每层需要9.48毫秒。Herald既可以在设计阶段用于进行协同优化(优化器),也可以在编译阶段用于进行层调度并报告预期的延迟和能耗(成本模型/调度器)。在我们的案例研究中,与每个评估设置下的最佳单片加速器相比,通过在HDA中部署两个互补式的子加速器,Herald所确定的经过协同优化的HDA在我们评估的工作负载和加速器中,平均在能耗延迟积(EDP)方面提高了56.0%,在延迟方面有46.82%的改善,在能耗方面有6.3%的优势。
New real time applications such as virtual reality (VR) with multiple sub-tasks (e.g., image segmentation, hand tracking, etc.) based on deep neural networks (DNNs) are emerging. Such applications rely on multiple DNNs for each sub-task with heterogeneity and require to meet target processing rates for each sub-task. Thus, the new compound DNN workload imposes new challenges to accelerator designs in two folds: (1) meeting target processing rate for each DNN for sub-tasks and (2) efficiently processing heterogeneous layers. As a solution, we explore heterogeneous DNN accelerators (HDAs). HDAs consist of multiple accelerator substrates (i.e., sub-accelerators) to support model parallelism that implements different mapping styles to provide adaptivity to heterogeneous DNN layers. However, HDA's performance and energy heavily depend on (1) how we partition hardware resources for sub-accelerators, and (2) how we schedule layers on the sub-accelerators. Therefore, we propose an HDA optimization framework, Herald, which performs co-optimization of hardware partitioning and layer schedluing. Herald exploits the simplicity of dependence graph of DNN layers to reduce the complexity of the problem, which on average requires 9.48 ms per layer on a laptop with i9-9880H processor and 16GB memory. Herald can be utilized both in design time to perform co-optimization (optimizer) and in compile time to perform layer scheduling and report expected latency and energy (cost model/scheduler). In our case studies, co-optimized HDAs with the best EDP Herald identified provided 56.0% EDP improvements with 46.82% latency and 6.3% energy benefits on average across workloads and accelerators we evaluate compared to the best monolithic accelerators for each evaluation setting by deploying two complementary-style sub-accelerators in an HDA.