Performance improvement of CUDA applications by reducing CPU-GPU data transfer overhead

Performance improvement of CUDA applications by reducing CPU-GPU data transfer overhead
复制标题

通过减少 CPU-GPU 数据传输开销提高 CUDA 应用程序的性能

DOI:
--
复制
发表时间:
2017
期刊:
International Conference Inventive Communication and Computational Technologies
影响因子:
--
通讯作者:
Niranjan N. Chiplunkar
Niranjan N. Chiplunkar
中科院分区:
--
文献类型:
--
作者:
N. V. Sunitha;K. Raju;Niranjan N. Chiplunkar

文献摘要

参考文献

被引文献

相似文献

在基于CPU-GPU的异构计算系统中,要由内核处理的输入数据驻留在主机存储器中。主机和设备内存地址空间不同。因此,设备不能直接访问主机存储器。在CUDA编程模型中,数据在主机存储器和设备存储器之间移动。这种数据传输是一项耗时的任务。通信开销可以通过重叠数据传输和内核执行来隐藏。CUDA流提供了一种重叠数据传输和内核执行的方法。在本文中,我们将探讨重叠的数据传输和内核执行的一些CUDA应用程序的整体执行时间的影响。结果表明,流支持的不同级别的并发的使用提高了CUDA应用程序的性能。
In a CPU-GPU based heterogeneous computing system, the input data to be processed by the kernel resides in the host memory. The host and the device memory address spaces are different. Therefore, the device can not directly access the host memory. In CUDA programming model, the data is moved between the host memory and the device memory. This data transfer is a time consuming task. The communication overhead can be hidden by overlapping the data transfer and the kernel execution. CUDA streams provide a means for overlapping data transfer and the kernel execution. In this paper we explore the effects of overlapping data transfer and the kernel execution on overall execution time of some CUDA applications. The results show that the usage of the different levels of concurrency supported by the streams enhances the performance of the CUDA applications.
DOI: --
发表时间: 2009
期刊: Scientific Reports
影响因子: 4.6
作者:
J. Xu
通讯作者: J. Xu