Photonic Reconfigurable Accelerators for Efficient Inference of CNNs With Mixed-Sized Tensors

Photonic Reconfigurable Accelerators for Efficient Inference of CNNs With Mixed-Sized Tensors
复制标题

DOI:
10.1109/tcad.2022.3197538
复制
发表时间:
2022-07
影响因子:
2.9
通讯作者:
Sairam Sri Vatsavai;Ishan G. Thakkar
Sairam Sri Vatsavai;Ishan G. Thakkar
中科院分区:
计算机科学3区
文献类型:
--
作者:
Sairam Sri Vatsavai;Ishan G. Thakkar

文献摘要

相似文献

基于光子微谐振器(MRR)的硬件加速器已被证明可以为处理深度卷积神经网络(CNN)提供破坏性的加速和能效改进。然而,以前基于MRR的CNN加速器无法为具有混合大小张量的CNN提供有效的适应性。这样的CNN的一个示例是深度可分离的CNN。在这种不灵活的加速器上使用混合大小的张量执行CNN的推理通常会导致硬件利用率低,这会降低加速器可实现的性能和能效。在本文中,我们提出了一种在基于MRR的CNN加速器中引入可重构性的新方法,以实现加速器硬件组件和使用硬件组件处理的CNN张量之间的大小兼容性的动态最大化。我们根据加速器中所用硬件组件的布局和相对位置,将现有工作中最先进的基于MRR的CNN加速器分为两类。然后,我们使用我们的方法在这两个类的加速器中引入可重构性,从而提高它们的并行性,有效映射不同大小,速度和整体能源效率的张量的灵活性。我们评估我们的可重构加速器对三个以前的作品面积比例的前景(所有加速器的硬件面积相等)。我们对四个现代CNN的推理进行的评估表明,与先前工作中基于MRR的加速器相比,我们设计的可重新配置的CNN加速器在每秒帧数(FPS)和FPS/W方面的改进高达1.8\times $和1.5\times $。
Photonic microring resonator (MRR)-based hardware accelerators have been shown to provide disruptive speedup and energy-efficiency improvements for processing deep convolutional neural networks (CNNs). However, previous MRR-based CNN accelerators fail to provide efficient adaptability for CNNs with mixed-sized tensors. One example of such CNNs is depthwise separable CNNs. Performing inferences of CNNs with mixed-sized tensors on such inflexible accelerators often leads to low hardware utilization, which diminishes the achievable performance and energy efficiency from the accelerators. In this article, we present a novel way of introducing reconfigurability in the MRR-based CNN accelerators, to enable dynamic maximization of the size compatibility between the accelerator hardware components and the CNN tensors that are processed using the hardware components. We classify the state-of-the-art MRR-based CNN accelerators from prior works into two categories, based on the layout and relative placements of the utilized hardware components in the accelerators. We then use our method to introduce reconfigurability in accelerators from these two classes, to consequently improve their parallelism, the flexibility of efficiently mapping tensors of different sizes, speed, and overall energy efficiency. We evaluate our reconfigurable accelerators against three prior works for the area proportionate outlook (equal hardware area for all accelerators). Our evaluation for the inference of four modern CNNs indicates that our designed reconfigurable CNN accelerators provide improvements of up to $1.8\times $ in frames-per-second (FPS) and up to $1.5\times $ in FPS/W, compared to an MRR-based accelerator from prior work.