A High-Performance and Energy-Efficient Photonic Architecture for Multi-DNN Acceleration

A High-Performance and Energy-Efficient Photonic Architecture for Multi-DNN Acceleration
复制标题

DOI:
10.1109/tpds.2023.3327535
复制
发表时间:
2024-01
影响因子:
5.3
通讯作者:
Yuan Li;A. Louri;Avinash Karanth
Yuan Li;A. Louri;Avinash Karanth
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yuan Li;A. Louri;Avinash Karanth

文献摘要

相似文献

大规模深度神经网络(DNN)加速器旨在促进不同深度神经网络的并发处理,这对互连结构提出了苛刻的挑战。这些挑战包括克服与系统扩展相关的性能下降和能量增加,同时还需要灵活性来支持动态分区和可适应的计算资源组织。然而,传统的基于金属的互连在可扩展性和灵活性方面经常面临固有的限制。在本文中,我们利用硅光子互连并采用算法架构协同设计方法来开发MDA,这是一种精心设计的DNN加速器,旨在实现各种DNN的高性能和节能并发处理。具体来说,MDA包括三个新的组成部分:1)资源分配算法,该算法根据并发dnn的计算需求和优先级将计算资源分配给并发dnn;2)数据流选择算法,确定每个DNN的片外和片上数据流,目标分别是最小化片外和片上存储器访问;3)一个灵活的硅光子网络,可以动态分割成子网络,每个子网络相互连接某个DNN的指定计算资源,同时适应所选择的片上数据流所决定的通信模式。仿真结果表明,所提出的MDA加速器优于其他最先进的多深度神经网络加速器,包括PREMA, AI-MT, Planaria和HDA。MDA加速器实现了3.6的加速提升,同时在能效、SLA满意率和公平性方面分别大幅提升了7.3倍、12.7倍和9.2倍。
Large-scale deep neural network (DNN) accelerators are poised to facilitate the concurrent processing of diverse DNNs, imposing demanding challenges on the interconnection fabric. These challenges encompass overcoming performance degradation and energy increase associated with system scaling while also necessitating flexibility to support dynamic partitioning and adaptable organization of compute resources. Nevertheless, conventional metallic-based interconnects frequently confront inherent limitations in scalability and flexibility. In this paper, we leverage silicon photonic interconnects and adopt an algorithm-architecture co-design approach to develop MDA, a DNN accelerator meticulously crafted to empower high-performance and energy-efficient concurrent processing of diverse DNNs. Specifically, MDA consists of three novel components: 1) a resource allocation algorithm that assigns compute resources to concurrent DNNs based on their computational demands and priorities; 2) a dataflow selection algorithm that determines off-chip and on-chip dataflows for each DNN, with the objectives of minimizing off-chip and on-chip memory accesses, respectively; 3) a flexible silicon photonic network that can be dynamically segmented into sub-networks, each interconnecting the assigned compute resources of a certain DNN while adapting to the communication patterns dictated by the selected on-chip dataflow. Simulation results show that the proposed MDA accelerator outperforms other state-of-the-art multi-DNN accelerators, including PREMA, AI-MT, Planaria, and HDA. MDA accelerator achieves a speedup of 3.6, accompanied by substantial improvements of 7.3×, 12.7×, and 9.2× in energy efficiency, service-level agreement (SLA) satisfaction rate, and fairness, respectively.