Heterogeneous Dataflow Accelerators for Multi-DNN Workloads

Heterogeneous Dataflow Accelerators for Multi-DNN Workloads
复制标题

DOI:
10.1109/hpca51647.2021.00016
复制
发表时间:
2020-12
期刊:
2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Hyoukjun Kwon;Liangzhen Lai;Michael Pellauer;T. Krishna;Yu-hsin Chen;V. Chandra
Hyoukjun Kwon;Liangzhen Lai;Michael Pellauer;T. Krishna;Yu-hsin Chen;V. Chandra
中科院分区:
其他
文献类型:
--
作者:
Hyoukjun Kwon;Liangzhen Lai;Michael Pellauer;T. Krishna;Yu-hsin Chen;V. Chandra

文献摘要

相似文献

增强现实和虚拟现实(AR/VR)等新兴AI应用利用多个深度神经网络(DNN)模型来执行各种子任务,如对象检测、图像分割、眼动跟踪、语音识别等。由于子任务的多样性,DNN模型内部和跨DNN模型的层在操作和形状上高度异构。对于在单个DNN加速器衬底上采用固定并行策略的固定并行加速器(FDA)来说,不同的层操作和形状是主要挑战,因为每个层偏好不同的并行(计算顺序和并行化)和瓦片大小。已经提出了可重构DNN加速器(RDA),以使其流适应不同的层,以应对这一挑战。然而,RDA的低灵活性是以昂贵的硬件结构(交换机、互连、控制器等)为代价的。并且需要每层重新配置,这引入了相当大的能量成本。或者,这项工作提出了一类新的加速器,异质低加速器(HDA),其部署多个加速器基底(即,子加速器),每个子加速器支持不同的子加速器。HDA比RDA具有更高的能效和更低的面积成本,比RDA具有更粗粒度的灵活性。为了利用这些好处,需要仔细优化跨子加速器的硬件资源分区和层执行调度。因此,我们还提出了先驱,一个框架,共同优化硬件分区和层调度。在一套AR/VR和MLPerf工作负载上使用Herald,我们发现了一种有前途的HDA架构Maelstrom,与最好的固定式低功耗加速器相比,它的延迟降低了65.3%,能耗降低了5.0%,与最先进的可重构DNN加速器(RDA)相比,能耗降低了22.0%,延迟增加了20.7%。结果表明,HDA是RDA的另一类帕累托最优加速器,具有能量优势,根据用例,它可能是比RDA更好的选择。
Emerging AI-enabled applications such as augmented and virtual reality (AR/VR) leverage multiple deep neural network (DNN) models for various sub-tasks such as object detection, image segmentation, eye-tracking, speech recognition, and so on. Because of the diversity of the sub-tasks, the layers within and across the DNN models are highly heterogeneous in operation and shape. Diverse layer operations and shapes are major challenges for a fixed dataflow accelerator (FDA) that employs a fixed dataflow strategy on a single DNN accelerator substrate since each layer prefers different dataflows (computation order and parallelization) and tile sizes. Reconfigurable DNN accelerators (RDAs) have been proposed to adapt their dataflows to diverse layers to address the challenge. However, the dataflow flexibility in RDAs is enabled at the cost of expensive hardware structures (switches, interconnects, controller, etc.) and requires per-layer reconfiguration, which introduces considerable energy costs. Alternatively, this work proposes a new class of accelerators, heterogeneous dataflow accelerators (HDAs), which deploy multiple accelerator substrates (i.e., sub-accelerators), each supporting a different dataflow. HDAs enable coarser-grained dataflow flexibility than RDAs with higher energy efficiency and lower area cost comparable to FDAs. To exploit such benefits, hardware resource partitioning across sub-accelerators and layer execution schedule need to be carefully optimized. Therefore, we also present Herald, a framework for co-optimizing hardware partitioning and layer scheduling. Using Herald on a suite of AR/VR and MLPerf workloads, we identify a promising HDA architecture, Maelstrom, which demonstrates 65.3% lower latency and 5.0% lower energy compared to the best fixed dataflow accelerators and 22.0% lower energy at the cost of 20.7% higher latency compared to a state-of-the-art reconfigurable DNN accelerator (RDA). The results suggest that HDA is an alternative class of Pareto-optimal accelerators to RDA with strength in energy, which can be a better choice than RDAs depending on the use cases.