Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient Training

Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient Training
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Anup Sarma;Sonali Singh;Huaipan Jiang;Rui Zhang;M. Kandemir;C. Das
Anup Sarma;Sonali Singh;Huaipan Jiang;Rui Zhang;M. Kandemir;C. Das
中科院分区:
其他
文献类型:
--
作者:
Anup Sarma;Sonali Singh;Huaipan Jiang;Rui Zhang;M. Kandemir;C. Das

文献摘要

相似文献

递归神经网络(RNN),更具体地说是它们的长短期记忆(LSTM)变体,已被广泛用作深度学习工具,用于处理文本和语音中基于序列的学习任务。这种LSTM应用程序的训练是计算密集型的,因为每个时间步重复的隐藏状态计算的递归性质。虽然深度神经网络中的稀疏性被广泛认为是减少训练和推理阶段计算时间的机会,但在LSTM RNN中使用非ReLU激活使得与神经元激活和梯度值相关的动态稀疏性的机会有限或不存在。在这项工作中,我们确定辍学引起的稀疏LSTM作为一个合适的模式,减少计算。Dropout是一种广泛使用的正则化机制,它在每次训练迭代期间随机丢弃计算的神经元值。我们建议结构辍学模式,通过丢弃同一组物理神经元内的一批,导致列(行)级隐藏状态稀疏,这是很好的服从计算减少在运行时的通用SIMD硬件以及脉动阵列。我们对三个代表性的NLP任务进行了实验:PTB数据集上的语言建模,使用IWALDE-EN和En-Vi数据集的基于OpenNMT的机器翻译,以及使用CoNLL-2003共享任务的命名实体识别序列标记。我们证明了我们提出的方法可以用于将基于丢弃的计算减少转化为减少的训练时间,改进范围从1.23倍到1.64倍,而不会牺牲目标度量。
Recurrent Neural Networks (RNNs), more specifically their Long Short-Term Memory (LSTM) variants, have been widely used as a deep learning tool for tackling sequence-based learning tasks in text and speech. Training of such LSTM applications is computationally intensive due to the recurrent nature of hidden state computation that repeats for each time step. While sparsity in Deep Neural Nets has been widely seen as an opportunity for reducing computation time in both training and inference phases, the usage of non-ReLU activation in LSTM RNNs renders the opportunities for such dynamic sparsity associated with neuron activation and gradient values to be limited or non-existent. In this work, we identify dropout induced sparsity for LSTMs as a suitable mode of computation reduction. Dropout is a widely used regularization mechanism, which randomly drops computed neuron values during each iteration of training. We propose to structure dropout patterns, by dropping out the same set of physical neurons within a batch, resulting in column (row) level hidden state sparsity, which are well amenable to computation reduction at run-time in general-purpose SIMD hardware as well as systolic arrays. We conduct our experiments for three representative NLP tasks: language modelling on the PTB dataset, OpenNMT based machine translation using the IWSLT De-En and En-Vi datasets, and named entity recognition sequence labelling using the CoNLL-2003 shared task. We demonstrate that our proposed approach can be used to translate dropout-based computation reduction into reduced training time, with improvement ranging from 1.23x to 1.64x, without sacrificing the target metric.