In Situ Stochastic Training of MTJ Crossbars With Machine Learning Algorithms

In Situ Stochastic Training of MTJ Crossbars With Machine Learning Algorithms
复制标题

DOI:
10.1145/3309880
复制
发表时间:
2019-03
期刊:
ACM Journal on Emerging Technologies in Computing Systems (JETC)
影响因子:
--
通讯作者:
Ankit Mondal;Ankur Srivastava-
Ankit Mondal;Ankur Srivastava-
中科院分区:
其他
文献类型:
--
作者:
Ankit Mondal;Ankur Srivastava-

文献摘要

被引文献

相似文献

基于磁隧道结(MTJ)的交叉棒由于具有高的器件密度、可扩展性和非易失性,在实现神经网络(NN)的权重方面引起了极大的兴趣。在MTJ中只存在两个稳定状态意味着在软件中获得最优二进制权重的高开销。本文说明交叉开关结构的内在并行性使其非常适合于现场训练,其中网络直接在硬件上示教。由于训练时间与网络的大小无关,因此它导致显著较小的训练开销,同时还绕过交叉开关中的交流电流路径的影响,并考虑到设备中的制造变化。我们展示了如何利用MTJ的随机切换特性来使用梯度下降算法来执行概率权重更新。我们描述了如何在实现NNS和受限Boltzmann机器的Crosbar上执行更新操作,并在它们上进行了模拟,以证明我们技术的有效性。结果表明,随机训练的MTJ-Crosbar前馈网络和深度信念网络的分类精度与软件训练的实值加权网络几乎相同,并且对设备变化具有免疫力。
Owing to high device density, scalability, and non-volatility, magnetic tunnel junction (MTJ)-based crossbars have garnered significant interest for implementing the weights of neural networks (NNs). The existence of only two stable states in MTJs implies a high overhead of obtaining optimal binary weights in software. This article illustrates that the inherent parallelism in the crossbar structure makes it highly appropriate for in situ training, wherein the network is taught directly on the hardware. It leads to significantly smaller training overhead as the training time is independent of the size of the network, while also circumventing the effects of alternate current paths in the crossbar and accounting for manufacturing variations in the device. We show how the stochastic switching characteristics of MTJs can be leveraged to perform probabilistic weight updates using the gradient descent algorithm. We describe how the update operations can be performed on crossbars implementing NNs and restricted Boltzmann machines, and perform simulations on them to demonstrate the effectiveness of our techniques. The results reveal that stochastically trained MTJ-crossbar feed-forward and deep belief nets achieve a classification accuracy nearly the same as that of real-valued weight networks trained in software and exhibit immunity to device variations.