GRNN: Low-Latency and Scalable RNN Inference on GPUs

GRNN: Low-Latency and Scalable RNN Inference on GPUs
复制标题

DOI:
10.1145/3302424.3303949
复制
发表时间:
2019-03
期刊:
Proceedings of the Fourteenth EuroSys Conference 2019
影响因子:
--
通讯作者:
Connor Holmes;Daniel Mawhirter;Yuxiong He;Feng Yan;Bo Wu
Connor Holmes;Daniel Mawhirter;Yuxiong He;Feng Yan;Bo Wu
中科院分区:
其他
文献类型:
--
作者:
Connor Holmes;Daniel Mawhirter;Yuxiong He;Feng Yan;Bo Wu

文献摘要

被引文献

相似文献

递归神经网络(RNNs)由于其在序列数据(如文本和语音信号)建模方面的有效性而受到广泛关注。然而,由于复杂的数据依赖关系和有限的并行性,目前gpu上用于rnn的推理库要么延迟高,要么可扩展性差,导致资源利用效率低下。因此,像微软和Facebook这样的公司使用cpu来服务RNN模型。这项工作从几个方面展示了gpu上现有RNN推理实现性能不理想的根本原因,包括数据重用性差、片上资源利用率低和高同步开销。我们系统地解决了这些问题,并开发了一个基于gpu的RNN推理库,称为GRNN,它提供了低延迟,高吞吐量和高效的资源利用。GRNN最大限度地减少了全局内存访问和同步开销,并通过新颖的数据重组、线程映射和性能建模技术平衡了片上资源的使用。通过对广泛的基准测试和实际应用进行评估,我们发现GRNN在延迟减少方面比最先进的CPU推理库高出17.5倍,比最先进的GPU推理库高出9倍。
Recurrent neural networks (RNNs) have gained significant attention due to their effectiveness in modeling sequential data, such as text and voice signal. However, due to the complex data dependencies and limited parallelism, current inference libraries for RNNs on GPUs produce either high latency or poor scalability, leading to inefficient resource utilization. Consequently, companies like Microsoft and Facebook use CPUs to serve RNN models. This work demonstrates the root causes of the unsatisfactory performance of existing implementations for RNN inference on GPUs from several aspects, including poor data reuse, low on-chip resource utilization, and high synchronization overhead. We systematically address these issues and develop a GPU-based RNN inference library, called GRNN, that provides low latency, high throughput, and efficient resource utilization. GRNN minimizes global memory accesses and synchronization overhead, as well as balancing on-chip resource usage through novel data reorganization, thread mapping, and performance modeling techniques. Evaluated on extensive benchmarking and real-world applications, we show that GRNN outperforms the state-of-the-art CPU inference library by up to 17.5X and state-of-the-art GPU inference libraries by up to 9X in terms of latency reduction.