Persistent RNNs: Stashing Weights on Chip
Persistent RNNs: Stashing Weights on Chip
复制标题
持久 RNN:将权重存储在芯片上
DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
S. Satheesh
中科院分区:
文献类型:
--
作者:
G. Diamos;Shubho Sengupta;Bryan Catanzaro;Mike Chrzanowski;Adam Coates;Erich Elsen;Jesse Engel;Awni Y. Hannun;S. Satheesh
This paper introduces a framework for mapping Recurrent Neural Network (RNN) architectures efficiently onto parallel processors such as GPUs. Key to our approach is the use of persistent computational kernels that exploit the processor’s memory hierarchy to reuse network weights over multiple timesteps. Using our framework, we show how it is possible to achieve substantially higher computational throughput at lower mini-batch sizes than direct implementations of RNNs based on matrix multiplications. Our initial implementation achieves 2.8 TFLOP/s at a mini-batch size of 4 on an NVIDIA TitanX GPU, which is about 45% of theoretical peak throughput, and is 30X faster than a standard RNN implementation based on optimized GEMM kernels at this batch size. Reducing the batch size from 64 to 4 per processor provides a 16x reduction in activation memory footprint, enables strong scaling to 16x more GPUs using data-parallelism, and allows us to efficiently explore end-to-end speech recognition models with up to 108 residual RNN layers.
DOI:
--
发表时间:
2015-12
期刊:
--
影响因子:
--
作者:
Dario Amodei;S. Ananthanarayanan;Rishita Anubhai;Jin Bai;Eric Battenberg;Carl Case;J. Casper;
通讯作者:
Dario Amodei;S. Ananthanarayanan;Rishita Anubhai;Jin Bai;Eric Battenberg;Carl Case;J. Casper;