Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural Networks

Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural Networks
复制标题

DOI:
10.1109/micro50266.2020.00062
复制
发表时间:
2020-10
期刊:
2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Soroush Ghodrati;Byung Hoon Ahn;J. Kim;Sean Kinzer;B. Yatham;N. Alla;Hardik Sharma;Mohammad Alian;Eiman Ebrahimi;N. Kim;C. Young;H. Esmaeilzadeh
Soroush Ghodrati;Byung Hoon Ahn;J. Kim;Sean Kinzer;B. Yatham;N. Alla;Hardik Sharma;Mohammad Alian;Eiman Ebrahimi;N. Kim;C. Young;H. Esmaeilzadeh
中科院分区:
其他
文献类型:
--
作者:
Soroush Ghodrati;Byung Hoon Ahn;J. Kim;Sean Kinzer;B. Yatham;N. Alla;Hardik Sharma;Mohammad Alian;Eiman Ebrahimi;N. Kim;C. Young;H. Esmaeilzadeh

文献摘要

被引文献

相似文献

深度神经网络 (DNN) 重振了依赖数据学习模式的现实世界应用,并渗透到不同的行业和市场。提供信息推理即服务 (INFaaS) 的云基础设施和加速器已成为行业中这一相当快速且侵入性转变的推动者。为此,大多数基于加速器的 INFaaS(Google 的 TPU [1]、NVIDIA T4 [2]、Microsoft Brainwave [3] 等)已成为许多现实应用程序的支柱。然而,随着对此类服务的需求不断增长,仅仅横向扩展加速器的数量在经济上并不划算。尽管多租户推动了数据中心的可扩展性,但由于对更高速度和效率的军备竞赛,它并不是设计 DNN 加速器的主要因素。本文着手通过一个新的维度:动态架构裂变来探索多租户的这一及时需求。为此,我们定义了 Planaria1,它可以在运行时动态分裂(分解)为多个更小但成熟的 DNN 引擎。这种微架构功能支持在同一硬件上在空间上共置多个 DNN 推理服务,从而提供同步多租户 DNN 加速。为了实现这种动态可重构性,我们首先设计了可破坏的全向脉动阵列,用于 DNN 加速,允许全向数据流。其次,它利用这种功能以及片上存储器、互连和计算资源的独特组织来实现基于脉动阵列的 DNN 加速器中的裂变。架构裂变及其相关的灵活性为任务调度提供了额外的自由度,甚至允许在服务器负载、DNN 拓扑和任务优先级方面打破加速器。因此,它可以同时并置 DNN,以提高利用率、吞吐量、QoS 和公平性。我们将所提出的设计与 PREMA [4] 进行比较,后者是最近的一项工作,通过跨多个任务时间复用 DNN 加速器来提供多租户。我们为两个加速器使用相同的频率、相同数量的计算和内存资源。结果显示,在(软、中、硬)QoS 要求、吞吐量(7.4×、7.2×、12.2×)、SLA 满意度(45%、15%、16%)和公平性(2.1×、2.3×、1.9×)方面都有显着优势。
Deep Neural Networks (DNNs) have reinvigorated real-world applications that rely on learning patterns of data and are permeating into different industries and markets. Cloud infrastructure and accelerators that offer INFerence-as-a-Service (INFaaS) have become the enabler of this rather quick and invasive shift in the industry. To that end, mostly accelerator-based INFaaS (Google’s TPU [1], NVIDIA T4 [2], Microsoft Brainwave [3], etc.) has become the backbone of many real-life applications. However, as the demand for such services grows, merely scaling-out the number of accelerators is not economically cost-effective. Although multi-tenancy has propelled datacenter scalability, it has not been a primary factor in designing DNN accelerators due to the arms race for higher speed and efficiency. This paper sets out to explore this timely requirement of multi-tenancy through a new dimension: dynamic architecture fission. To that end, we define Planaria1 that can dynamically fission (break) into multiple smaller yet full-fledged DNN engines at runtime. This microarchitectural capability enables spatially co-locating multiple DNN inference services on the same hardware, offering simultaneous multi-tenant DNN acceleration. To realize this dynamic reconfigurability, we first devise breakable omni-directional systolic arrays for DNN acceleration that allows omni-directional flow of data. Second, it uses this capability and a unique organization of on-chip memory, interconnection, and compute resources to enable fission in systolic array based DNN accelerators. Architecture fission and its associated flexibility enables an extra degree of freedom for task scheduling, that even allows breaking the accelerator with regard to the server load, DNN topology, and task priority. As such, it can simultaneously co-locate DNNs to enhance utilization, throughput, QoS, and fairness. We compare the proposed design to PREMA [4], a recent effort that offers multi-tenancy by time-multiplexing the DNN accelerator across multiple tasks. We use the same frequency, the same amount of compute and memory resources for both accelerators. The results show significant benefits with (soft, medium, hard) QoS requirements, in throughput (7.4×, 7.2×, 12.2×), SLA satisfaction rate (45%, 15%, 16%), and fairness (2.1×, 2.3×, 1.9×).