Data Transfer Matters for GPU Computing

Data Transfer Matters for GPU Computing
复制标题

DOI:
10.1109/icpads.2013.47
复制
发表时间:
2013-12
期刊:
2013 International Conference on Parallel and Distributed Systems
影响因子:
--
通讯作者:
Yusuke Fujii;Takuya Azumi;N. Nishio;S. Kato;M. Edahiro
Yusuke Fujii;Takuya Azumi;N. Nishio;S. Kato;M. Edahiro
中科院分区:
其他
文献类型:
--
作者:
Yusuke Fujii;Takuya Azumi;N. Nishio;S. Kato;M. Edahiro

文献摘要

被引文献

相似文献

图形处理单元(GPU)包含众核计算设备,其中大规模并行计算线程从CPU卸载。GPU计算的这种异构性质引起了重要的数据传输问题,特别是针对延迟关键的实时系统。然而,即使是与GPU计算相关的数据传输的基本特征在文献中也没有得到很好的研究。在本文中,我们调查和表征目前可实现的数据传输方法的尖端GPU技术。我们使用开源软件来实现这些方法,以比较它们在现实系统中的性能和延迟。我们的实验结果表明,硬件辅助的直接内存访问(DMA)和I/O读写访问方法通常是最有效的,而片上微控制器内的GPU是有用的,在减少并发多个数据流的数据传输延迟。我们还公开了CPU优先级可以保护GPU数据传输的性能。
Graphics processing units (GPUs) embrace many-core compute devices where massively parallel compute threads are offloaded from CPUs. This heterogeneous nature of GPU computing raises non-trivial data transfer problems especially against latency-critical real-time systems. However even the basic characteristics of data transfers associated with GPU computing are not well studied in the literature. In this paper, we investigate and characterize currently-achievable data transfer methods of cutting-edge GPU technology. We implement these methods using open-source software to compare their performance and latency for real-world systems. Our experimental results show that the hardware-assisted direct memory access (DMA) and the I/O read-and-write access methods are usually the most effective, while on-chip micro controllers inside the GPU are useful in terms of reducing the data transfer latency for concurrent multiple data streams. We also disclose that CPU priorities can protect the performance of GPU data transfers.