The Random Address Shift to Reduce the Memory Access Congestion on the Discrete Memory Machine

The Random Address Shift to Reduce the Memory Access Congestion on the Discrete Memory Machine
复制标题

随机地址移位减少离散存储器上的存储器访问拥塞

DOI:
10.1109/candar.2013.21
复制
发表时间:
2013
期刊:
2013 First International Symposium on Computing and Networking
影响因子:
--
通讯作者:
Yasuaki Ito
Yasuaki Ito
中科院分区:
--
文献类型:
--
作者:
K. Nakano;Susumu Matsumae;Yasuaki Ito

文献摘要

被引文献

相似文献

离散内存机(DMM)是一种理论上的并行计算模型,它捕获了支持cuda的gpu上流多处理器的内存访问的本质。DMM有w个组成共享内存的内存库,并且有w个线程试图同时访问它们。然而,预定到同一内存库的内存访问请求是顺序处理的。因此,开发有效的算法来减少内存访问拥塞,即到达同一银行的内存访问请求的最大数量是非常重要的。内存访问拥塞的值在1到w之间。本文的主要贡献是提出了一种新的算法技术,称为随机地址移位,可以减少内存访问拥塞。我们表明,对于任何内存访问请求(包括w个线程的恶意请求),内存访问拥塞的预期值为O(log w/log log w)。仿真结果表明,w=32线程的预期拥塞率仅为3.436。由于针对同一银行的恶意内存访问请求占用拥塞32,我们的随机地址移位技术大大减少了内存访问拥塞。我们将随机地址移位技术应用到矩阵转置算法中。在GeForce GTX Titan上的实验结果表明,随机地址移位技术是实用的,可以将简单的矩阵转置算法的速度提高5倍。
The Discrete Memory Machine (DMM) is a theoretical parallel computing model that captures the essence of memory access of the streaming multiprocessor on CUDA-enabled GPUs. The DMM has w memory banks that constitute a shared memory, and w threads in a warp try to access them at the same time. However, memory access requests destined for the same memory bank are processed sequentially. Hence, it is very important for developing efficient algorithms to reduce the memory access congestion, the maximum number of memory access requests destined for the same bank. The memory access congestion takes value between 1 and w. The main contribution of this paper is to present a novel algorithmic technique called the random address shift that reduces the memory access congestion. We show that the memory access congestion is expected O(log w/log log w) for any memory access requests including malicious ones by a warp of w threads. The simulation results show that the expected congestion for w=32 threads is only 3.436. Since the malicious memory access requests destined for the same bank take congestion 32, our random address shift technique substantially reduces the memory access congestion. We have applied the random address shift technique to matrix transpose algorithms. The experimental results on GeForce GTX Titan show that the random address shift technique is practical and can accelerate the straightforward matrix transpose algorithms by a factor of 5.