VLSI Architectures for the Restricted Boltzmann Machine

VLSI Architectures for the Restricted Boltzmann Machine
复制标题

受限玻尔兹曼机的 VLSI 架构

DOI:
10.1145/3007193
复制
发表时间:
2017
期刊:
ACM Journal on Emerging Technologies in Computing Systems (JETC)
影响因子:
--
通讯作者:
K. Parhi
K. Parhi
中科院分区:
--
文献类型:
--
作者:
Bo Yuan;K. Parhi

文献摘要

被引文献

相似文献

神经网络(NN)系统广泛应用于从计算机视觉到语音识别的许多重要应用中。迄今为止,大多数神经网络系统都是由 CPU 或 GPU 等通用处理单元进行处理。然而,随着数据集和网络规模的迅速增加,原始软件实现的训练时间过长。为了克服这个问题,需要专门的硬件加速器来设计高速神经网络系统。本文提出了一种高效的受限玻尔兹曼机(RBM)硬件架构,它是神经网络系统的一个重要类别。在硬件层面执行各种优化方法来提高训练速度。使用尽快和重叠调度方法来减少延迟。结果表明,与扁平化设计相比,所提出的 RBM 架构可以实现训练时间减少 50%。此外,还采用了即时计算方案,将二进制和随机状态的存储需求减少了数百倍。然后,基于所提出的方法,为 MNIST 手写数字识别数据集开发了 784-2252 RBM 设计示例。分析表明,与基于CPU/GPU的解决方案相比,RBM的VLSI设计在训练速度和能源效率方面取得了显着提高。
Neural network (NN) systems are widely used in many important applications ranging from computer vision to speech recognition. To date, most NN systems are processed by general processing units like CPUs or GPUs. However, as the sizes of dataset and network rapidly increase, the original software implementations suffer from long training time. To overcome this problem, specialized hardware accelerators are needed to design high-speed NN systems. This article presents an efficient hardware architecture of restricted Boltzmann machine (RBM) that is an important category of NN systems. Various optimization approaches at the hardware level are performed to improve the training speed. As-soon-as-possible and overlapped-scheduling approaches are used to reduce the latency. It is shown that, compared with the flat design, the proposed RBM architecture can achieve 50% reduction in training time. In addition, an on-the-fly computation scheme is also used to reduce the storage requirement of binary and stochastic states by several hundreds of times. Then, based on the proposed approach, a 784-2252 RBM design example is developed for MNIST handwritten digit recognition dataset. Analysis shows that the VLSI design of RBM achieves significant improvement in training speed and energy efficiency as compared to CPU/GPU-based solution.