Performance improvement of CUDA applications by reducing CPU-GPU data transfer overhead
Performance improvement of CUDA applications by reducing CPU-GPU data transfer overhead
复制标题
通过减少 CPU-GPU 数据传输开销提高 CUDA 应用程序的性能
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Niranjan N. Chiplunkar
中科院分区:
文献类型:
--
作者:
N. V. Sunitha;K. Raju;Niranjan N. Chiplunkar
In a CPU-GPU based heterogeneous computing system, the input data to be processed by the kernel resides in the host memory. The host and the device memory address spaces are different. Therefore, the device can not directly access the host memory. In CUDA programming model, the data is moved between the host memory and the device memory. This data transfer is a time consuming task. The communication overhead can be hidden by overlapping the data transfer and the kernel execution. CUDA streams provide a means for overlapping data transfer and the kernel execution. In this paper we explore the effects of overlapping data transfer and the kernel execution on overall execution time of some CUDA applications. The results show that the usage of the different levels of concurrency supported by the streams enhances the performance of the CUDA applications.
影响因子:
4.6
作者:
J. Xu
通讯作者:
J. Xu