Device-Circuit-Architecture Co-Exploration for Computing-in-Memory Neural Accelerators

Device-Circuit-Architecture Co-Exploration for Computing-in-Memory Neural Accelerators
复制标题

DOI:
10.1109/tc.2020.2991575
复制
发表时间:
2019-10
影响因子:
3.7
通讯作者:
Weiwen Jiang;Qiuwen Lou;Zheyu Yan;Lei Yang;J. Hu;X. Hu;Yiyu Shi
Weiwen Jiang;Qiuwen Lou;Zheyu Yan;Lei Yang;J. Hu;X. Hu;Yiyu Shi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Weiwen Jiang;Qiuwen Lou;Zheyu Yan;Lei Yang;J. Hu;X. Hu;Yiyu Shi

文献摘要

相似文献

神经架构和硬件设计的共同探索是有前途的,因为它能够同时优化网络精度和硬件效率。然而,用于共同探索的最先进的神经架构搜索算法专用于传统的冯-诺依曼计算架构,其性能受到众所周知的存储器墙的严重限制。在本文中,我们首次将内存计算架构(可以轻松超越内存墙)与神经架构搜索相互作用,旨在找到具有高网络精度和最大硬件效率的最有效神经架构。这种新颖的组合为提高性能提供了机会,但也带来了一系列挑战:优化空间跨越从器件类型和电路拓扑到神经架构的多个设计层;器件变化的存在可能会大大降低神经网络的性能。为了解决这些挑战,我们提出了一个跨层的探索框架,即NACIM,它共同探索设备,电路和架构设计空间,并考虑到设备的变化,以找到最强大的神经架构,再加上最有效的硬件设计。实验结果表明,NACIM可以找到鲁棒的神经网络,在存在设备变化的情况下,准确率损失为0.45%,而不考虑变化的最先进的NAS的准确率损失为76.44%;此外,NACIM实现了高达16.3 TOPs/W的能量效率,比最先进的NAS高3.17倍。
Co-exploration of neural architectures and hardware design is promising due to its capability to simultaneously optimize network accuracy and hardware efficiency. However, state-of-the-art neural architecture search algorithms for the co-exploration are dedicated for the conventional von-Neumann computing architecture, whose performance is heavily limited by the well-known memory wall. In this article, we are the first to bring the computing-in-memory architecture, which can easily transcend the memory wall, to interplay with the neural architecture search, aiming to find the most efficient neural architectures with high network accuracy and maximized hardware efficiency. Such a novel combination makes opportunities to boost performance, but also brings a bunch of challenges: The optimization space spans across multiple design layers from device type and circuit topology to neural architecture; and the presence of device variation may drastically degrade the neural network performance. To address these challenges, we propose a cross-layer exploration framework, namely NACIM, which jointly explores device, circuit and architecture design space and takes device variation into consideration to find the most robust neural architectures, coupled with the most efficient hardware design. Experimental results demonstrate that NACIM can find the robust neural network with 0.45 percent accuracy loss in the presence of device variation, compared with a 76.44 percent loss from the state-of-the-art NAS without consideration of variation; in addition, NACIM achieves an energy efficiency up to 16.3 TOPs/W, 3.17× higher than the state-of-the-art NAS.