Scalable and Sustainable Deep Learning via Randomized Hashing

Scalable and Sustainable Deep Learning via Randomized Hashing
复制标题

DOI:
10.1145/3097983.3098035
复制
发表时间:
2016-02
期刊:
Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Ryan Spring;Anshumali Shrivastava
Ryan Spring;Anshumali Shrivastava
中科院分区:
其他
文献类型:
--
作者:
Ryan Spring;Anshumali Shrivastava

文献摘要

被引文献

相似文献

目前的深度学习体系结构正在增长,以便从复杂的数据集中学习。这些体系结构需要巨大的矩阵乘法操作来训练数百万个参数。相反,还有另一种增长的趋势是将深度学习带入低功率,嵌入式设备。从计算和能量的角度来看,与深层网络的培训和测试相关的矩阵操作非常昂贵。我们提出了一种基于哈希的新技术,可以大大减少训练和测试神经网络所需的计算量。我们的方法结合了两个最近的想法,即自适应辍学和随机散列,以进行最大的内部产品搜索(MIPS),以有效地选择具有最高激活的节点。我们的深度学习算法通过在较少的节点上操作,从而降低了前向和向后传播步骤的总体计算成本。因此,我们的算法仅使用总乘法的5%,同时平均保持原始模型准确性的1%。提议的基于哈希的后传播的独特属性是,更新总是很稀疏。由于梯度更新,我们的算法非常适合异步,并行训练,随着核心数量的增加而导致近线性加速。我们通过在几个数据集上进行了严格的实验评估来证明我们提出的算法的可伸缩性和可持续性(能效)。
Current deep learning architectures are growing larger in order to learn from complex datasets. These architectures require giant matrix multiplication operations to train millions of parameters. Conversely, there is another growing trend to bring deep learning to low-power, embedded devices. The matrix operations, associated with the training and testing of deep networks, are very expensive from a computational and energy standpoint. We present a novel hashing-based technique to drastically reduce the amount of computation needed to train and test neural networks. Our approach combines two recent ideas, Adaptive Dropout and Randomized Hashing for Maximum Inner Product Search (MIPS), to select the nodes with the highest activations efficiently. Our new algorithm for deep learning reduces the overall computational cost of the forward and backward propagation steps by operating on significantly fewer nodes. As a consequence, our algorithm uses only 5% of the total multiplications, while keeping within 1% of the accuracy of the original model on average. A unique property of the proposed hashing-based back-propagation is that the updates are always sparse. Due to the sparse gradient updates, our algorithm is ideally suited for asynchronous, parallel training, leading to near-linear speedup, as the number of cores increases. We demonstrate the scalability and sustainability (energy efficiency) of our proposed algorithm via rigorous experimental evaluations on several datasets.