SPRINT: A High-Performance, Energy-Efficient, and Scalable Chiplet-based Accelerator with Photonic Interconnects for CNN Inference

SPRINT: A High-Performance, Energy-Efficient, and Scalable Chiplet-based Accelerator with Photonic Interconnects for CNN Inference
复制标题

DOI:
10.1109/tpds.2021.3139015
复制
发表时间:
2021
影响因子:
5.3
通讯作者:
Yuan Li;A. Louri;Avinash Karanth
Yuan Li;A. Louri;Avinash Karanth
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yuan Li;A. Louri;Avinash Karanth

文献摘要

相似文献

基于Chiplet的卷积神经网络(CNN)加速器已经成为一种很有前途的解决方案,可以为CNN推理提供大量的处理能力和片上存储容量。这些加速器的性能常常受到芯片间金属互连的限制。新兴的技术,如光子互连,可以克服金属互连的限制,因为它具有高带宽密度和与距离无关的延迟等几个优越特性。然而,在基于芯片的CNN加速器中实现光子互连是具有挑战性的,需要网络架构优化和CNN数据流定制的共同努力。在本文中,我们提出了一种基于芯片的CNN加速器Sprint,它由一个全局缓冲区和几个加速器芯片组成。Sprint推出了两种新颖的设计:(1)光子芯片间网络,可通过波长分配和波导重新配置适应CNN推理中的特定通信模式;(2)CNN数据流,可利用光子互连的广播能力,同时最大限度地减少昂贵的电-光和光-电信号转换。使用多个CNN模型的模拟表明,与其他最先进的基于金属或光子互连的基于芯片的体系结构相比,Sprint的执行时间和能耗分别减少了76%和68%。
Chiplet-based convolution neural network (CNN) accelerators have emerged as a promising solution to provide substantial processing power and on-chip memory capacity for CNN inference. The performance of these accelerators is often limited by inter-chiplet metallic interconnects. Emerging technologies such as photonic interconnects can overcome the limitations of metallic interconnects due to several superior properties including high bandwidth density and distance-independent latency. However, implementing photonic interconnects in chiplet-based CNN accelerators is challenging and requires combined effort of network architectural optimization and CNN dataflow customization. In this paper, we propose SPRINT, a chiplet-based CNN accelerator that consists of a global buffer and several accelerator chiplets. SPRINT introduces two novel designs: (1) a photonic inter-chiplet network that can adapt to specific communication patterns in CNN inference through wavelength allocation and waveguide reconfiguration, and (2) a CNN dataflow that can leverage the broadcasting capability of photonic interconnects while minimizing the costly electrical-to-optical and optical-to-electrical signal conversions. Simulations using multiple CNN models show that SPRINT achieves up to 76% and 68% reduction in execution time and energy consumption, respectively, as compared to other state-of-the-art chiplet-based architectures with either metallic or photonic interconnects.