SMAUG

SMAUG
复制标题

DOI:
10.1145/3424669
复制
发表时间:
2019-12
期刊:
ACM Transactions on Architecture and Code Optimization (TACO)
影响因子:
--
通讯作者:
S. Xi;Yuan Yao;K. Bhardwaj;P. Whatmough;Gu-Yeon Wei;D. Brooks
S. Xi;Yuan Yao;K. Bhardwaj;P. Whatmough;Gu-Yeon Wei;D. Brooks
中科院分区:
其他
文献类型:
--
作者:
S. Xi;Yuan Yao;K. Bhardwaj;P. Whatmough;Gu-Yeon Wei;D. Brooks

文献摘要

被引文献

相似文献

近年来,深度神经网络在硬件加速方面取得了巨大的进步。然而,大多数研究都集中在优化加速器微架构上,以便在每层的基础上获得更高的性能和能效。我们发现,对于整体单批推理延迟,加速器可能只占25-40%,其余的用于数据移动和深度学习软件框架。到目前为止,在早期设计阶段(在RTL可用之前)研究端到端DNN性能是非常困难的,因为没有现有的DNN框架支持端到端仿真,并且易于定制硬件加速器集成。为了解决研究基础设施方面的这一差距,我们提出了SMAUG,这是第一个专为模拟端到端深度学习应用而构建的深度神经网络框架。SMAUG为研究人员提供了广泛的能力来评估DNN工作负载,从不同的网络拓扑到简单的加速器建模和SoC集成。为了展示SMAUG的功能和价值,我们展示了案例研究,展示了我们如何优化整体性能和能源效率,在基线系统上加速1.8×-5×,而不改变加速器微架构的任何部分,以及SMAUG如何为摄像头驱动的深度学习管道调整SoC。
In recent years, there has been tremendous advances in hardware acceleration of deep neural networks. However, most of the research has focused on optimizing accelerator microarchitecture for higher performance and energy efficiency on a per-layer basis. We find that for overall single-batch inference latency, the accelerator may only make up 25–40%, with the rest spent on data movement and in the deep learning software framework. Thus far, it has been very difficult to study end-to-end DNN performance during early stage design (before RTL is available), because there are no existing DNN frameworks that support end-to-end simulation with easy custom hardware accelerator integration. To address this gap in research infrastructure, we present SMAUG, the first DNN framework that is purpose-built for simulation of end-to-end deep learning applications. SMAUG offers researchers a wide range of capabilities for evaluating DNN workloads, from diverse network topologies to easy accelerator modeling and SoC integration. To demonstrate the power and value of SMAUG, we present case studies that show how we can optimize overall performance and energy efficiency for up to 1.8×–5× speedup over a baseline system, without changing any part of the accelerator microarchitecture, as well as show how SMAUG can tune an SoC for a camera-powered deep learning pipeline.