Hardware transactional memory for GPU architectures

Hardware transactional memory for GPU architectures
复制标题

DOI:
10.1145/2155620.2155655
复制
发表时间:
2011-12
期刊:
2011 44th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Wilson W. L. Fung;Inderpreet Singh;Andrew Brownsword;Tor M. Aamodt
Wilson W. L. Fung;Inderpreet Singh;Andrew Brownsword;Tor M. Aamodt
中科院分区:
其他
文献类型:
--
作者:
Wilson W. L. Fung;Inderpreet Singh;Andrew Brownsword;Tor M. Aamodt

文献摘要

被引文献

相似文献

图形处理器单元 (GPU) 旨在有效利用线程级并行性 (TLP),在相对较小的一组单指令多线程 (SIMT) 核心上复用执行 1000 个并发线程,以隐藏各种长延迟操作。虽然 CUDA 块/OpenCL 工作组中的线程可以通过内核内暂存存储器进行有效通信,但不同块中的线程只能通过全局内存访问进行通信。希望利用此类通信的程序员必须考虑当多个线程修改同一内存位置时可能发生的数据争用。最近的 GPU 通过单个 32 位/64 位字的原子操作提供了一种块间通信形式。尽管可以从这些原子操作构造细粒度锁,但是使用锁的同步很容易出现死锁。在本文中,我们建议通过扩展 GPU 以支持事务内存 (TM) 来解决这些问题。主要挑战包括支持数千个并发事务并并行提交非冲突事务。我们提出了 KILO TM,这是一种新颖的 GPU 硬件 TM 设计,可扩展到 1000 个并发事务。它无需依赖缓存一致性硬件,而是使用字级、基于值的冲突检测来避免广播通信并减少片上存储开销。它使用新颖的布隆过滤器组织进行推测验证,以提高事务提交并行性。对于一组 TM 增强型 GPU 应用程序,KILO TM 捕获了 59% 的细粒度锁定性能,并且平均比串行执行所有事务快 128 倍,估计硬件面积开销为商用 GPU 的 0.5%。
Graphics processor units (GPUs) are designed to efficiently exploit thread level parallelism (TLP), multiplexing execution of 1000s of concurrent threads on a relatively smaller set of single-instruction, multiple-thread (SIMT) cores to hide various long latency operations. While threads within a CUDA block/OpenCL workgroup can communicate efficiently through an intra-core scratchpad memory, threads in different blocks can only communicate via global memory accesses. Programmers wishing to exploit such communication have to consider data-races that may occur when multiple threads modify the same memory location. Recent GPUs provide a form of inter-block communication through atomic operations for single 32-bit/64-bit words. Although fine-grained locks can be constructed from these atomic operations, synchronization using locks is prone to deadlock. In this paper, we propose to solve these problems by extending GPUs to support transactional memory (TM). Major challenges include supporting 1000s of concurrent transactions and committing non-conflicting transactions in parallel. We propose KILO TM, a novel hardware TM design for GPUs that scales to 1000s of concurrent transactions. Without cache coherency hardware to depend on, it uses word-level, value-based conflict detection to avoid broadcast communication and reduce on-chip storage overhead. It employs speculative validation using a novel bloomfilter organization to increase transaction commit parallelism. For a set of TM-enhanced GPU applications, KILO TM captures 59% of the performance of fine-grained locking, and is on average 128×faster than executing all transactions serially, for an estimated hardware area overhead of 0.5% of a commercial GPU.