ATT: A Fault-Tolerant ReRAM Accelerator for Attention-based Neural Networks

ATT: A Fault-Tolerant ReRAM Accelerator for Attention-based Neural Networks
复制标题

DOI:
10.1109/iccd50377.2020.00047
复制
发表时间:
2020-10
期刊:
2020 IEEE 38th International Conference on Computer Design (ICCD)
影响因子:
--
通讯作者:
Haoqiang Guo;Lu Peng;Jian Zhang-;Qing Chen;Travis LeCompte
Haoqiang Guo;Lu Peng;Jian Zhang-;Qing Chen;Travis LeCompte
中科院分区:
其他
文献类型:
--
作者:
Haoqiang Guo;Lu Peng;Jian Zhang-;Qing Chen;Travis LeCompte

文献摘要

被引文献

相似文献

纵横制阻性随机存储器在深度学习加速器设计中得到了广泛的应用,因为它在很大程度上消除了存储器和处理单元之间的重量移动。高密度存储和低泄漏功耗使其非常适合EDGE/IoT设备。然而,现有的传统神经网络的ReRAM设计不能支持基于注意力的神经网络,基于注意力的神经网络是由编码器和解码器堆叠而不是卷积层或全连通层。除了传统神经网络中的矩阵-矩阵乘法外,编解码器还包括注意机制、层归一化和高斯误差线性单元。这些新特性使得数据流比卷积层的数据流复杂得多。当映射严重降低计算精度的权重时,错误的ReRAM设备是额外的障碍。现有的硬件冗余策略不了解应用程序特性,通常会导致设计效率低下。在这项工作中,我们分析了这些基于注意力的神经网络的数据流,并提出了一种基于ReRAM的加速器,该加速器具有专用的流水线设计。当考虑交叉开关中存在硬故障的单元时,我们进一步提出了一种非均匀冗余策略NuXG,通过降低冗余度来满足精度要求并节省能源消耗。最后,我们对结果进行了评估,结果表明,在基于注意力的神经网络中,所提出的冗余方案在功率效率和吞吐量方面都比现有的冗余方案提高了两倍以上。此外,它的性能也大大超过了NVIDIA图形处理器。
Crossbar-based resistive RAM has been widely used in deep learning accelerator designs because it largely eliminates weight movement between memory and processing units. The high-density storage and low leakage power make it a good fit for edge/IoT devices. However, existing ReRAM designs for traditional neural networks cannot support Attention-based Neural Networks, which are stacked with encoders and decoders instead of convolutional layers or fully connected layers. In addition to matrix-matrix multiplications in traditional neural networks, an encoder or a decoder also includes the attention mechanism, the layer normalization and the gaussian error linear unit. These new characteristics make the data flow far more complicated than that of a convolutional layer. Faulty ReRAM devices are additional obstacles when mapping weights that severely degrade computation accuracy. Existing hardware redundancy strategies that are unaware of application characteristics usually result in inefficient designs. In this work, we analyze the data flow of these attention-based neural networks and propose a ReRAM-based accelerator with a dedicated pipeline design for Attention-based Neural Networks. When considering cells with hard faults in crossbars, we further propose NuXG, a non-uniform redundancy strategy, to meet accuracy requirements and save energy consumption by decreasing the redundancy ratio. Finally, we evaluate results and demonstrate that the proposed can achieve more than two times improved performance over existing redundancy schemes in both power efficiency and throughput for Attention-based Neural Networks. Moreover, it also significantly outperforms an NVIDIA GPU.