HiMA: A Fast and Scalable History-based Memory Access Engine for Differentiable Neural Computer

HiMA: A Fast and Scalable History-based Memory Access Engine for Differentiable Neural Computer
复制标题

DOI:
10.1145/3466752.3480052
复制
发表时间:
2021-10
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Yaoyu Tao;Zhengya Zhang
Yaoyu Tao;Zhengya Zhang
中科院分区:
其他
文献类型:
--
作者:
Yaoyu Tao;Zhengya Zhang

文献摘要

被引文献

相似文献

在外部存储器的帮助下,记忆增强神经网络(MANNs)在许多任务中提供了更好的推理性能。最近开发的可微分神经计算机(DNC)是一种MANN,已被证明在表示复杂数据结构和学习长期依赖关系方面表现优异。DNC的高性能来源于新的基于历史的注意机制以及之前使用的基于内容的注意机制。基于历史的机制需要各种新的计算原语和状态存储器,而现有的神经网络或MANN加速器不支持这些原语和状态存储器。我们提出了HiMA,这是一个基于历史的内存访问引擎,其内存分布在各个块中。HiMA采用多模片上网络(NoC)来减少通信延迟并提高可扩展性。采用最优子矩阵内存分区策略来减少NoC流量;两阶段使用排序方法利用分布式块来提高计算速度。为了使HiMA从根本上可扩展,我们创建了一个分布式版本的DNC,称为DNC- d,允许几乎所有的内存操作应用于具有可训练加权求和的本地内存,以产生全局内存输出。为了进一步提高硬件效率,提出了usage skimming和softmax两种近似技术。HiMA原型是在RTL中创建的,并在40nm技术中合成。通过仿真,HiMA运行DNC和DNC- d的速度分别比MANN加速器高6.47倍和39.1倍,面积效率分别提高22.8倍和164.3倍,能量效率分别提高6.1倍和61.2倍。与Nvidia 3080Ti GPU相比,HiMA在运行DNC和DNC- d时的加速分别高达437倍和2646倍。
Memory-augmented neural networks (MANNs) provide better inference performance in many tasks with the help of an external memory. The recently developed differentiable neural computer (DNC) is a MANN that has been shown to outperform in representing complicated data structures and learning long-term dependencies. DNC’s higher performance is derived from new history-based attention mechanisms in addition to the previously used content-based attention mechanisms. History-based mechanisms require a variety of new compute primitives and state memories, which are not supported by existing neural network (NN) or MANN accelerators. We present HiMA, a tiled, history-based memory access engine with distributed memories in tiles. HiMA incorporates a multi-mode network-on-chip (NoC) to reduce the communication latency and improve scalability. An optimal submatrix-wise memory partition strategy is applied to reduce the amount of NoC traffic; and a two-stage usage sort method leverages distributed tiles to improve computation speed. To make HiMA fundamentally scalable, we create a distributed version of DNC called DNC-D to allow almost all memory operations to be applied to local memories with trainable weighted summation to produce the global memory output. Two approximation techniques, usage skimming and softmax approximation, are proposed to further enhance hardware efficiency. HiMA prototypes are created in RTL and synthesized in a 40nm technology. By simulations, HiMA running DNC and DNC-D demonstrates 6.47 × and 39.1 × higher speed, 22.8 × and 164.3 × better area efficiency, and 6.1 × and 61.2 × better energy efficiency over the state-of-the-art MANN accelerator. Compared to an Nvidia 3080Ti GPU, HiMA demonstrates speedup by up to 437 × and 2,646 × when running DNC and DNC-D, respectively.