Throughput-oriented GPU memory allocation

Throughput-oriented GPU memory allocation
复制标题

面向吞吐量的 GPU 内存分配

DOI:
--
复制
发表时间:
2019
期刊:
ACM SIGPLAN Symposium on Principles & Practice of Parallel Programming
影响因子:
--
通讯作者:
M. Garland
M. Garland
中科院分区:
--
文献类型:
--
作者:
Isaac Gelado;M. Garland

文献摘要

被引文献

相似文献

面向吞吐量的架构(如GPU)可以支持比多核架构多三个数量级的并发线程。这种级别的并发会使典型的同步原语(例如,互斥体)超出其可伸缩性限制,从而在使用它们的模块(如内存分配器)中造成严重的性能瓶颈。在本文中,我们开发了并发编程技术和同步原语,以支持动态内存分配器,这些技术和同步原语对于使用非常高级别的并发是有效的。我们将资源分配描述为一个两阶段的过程,将可用资源的数量与对可用资源本身的跟踪分离。为了方便记账阶段,我们引入了一种新的批量信号量抽象,通过优化线程同时对信号量进行操作的情况来扩展传统的信号量语义。我们还类似地设计了新的集合同步原语,使协作线程组能够一起进入临界区。最后,我们展示了将延迟回收委托给已被阻塞的线程极大地提高了效率。使用所有这些技术,我们面向吞吐量的内存分配器在现代GPU上提供高分配率和低内存碎片。我们的实验表明,它的分配率平均是CUDA 9工具包中相应实现的16.56倍。
Throughput-oriented architectures, such as GPUs, can sustain three orders of magnitude more concurrent threads than multicore architectures. This level of concurrency pushes typical synchronization primitives (e.g., mutexes) over their scalability limits, creating significant performance bottlenecks in modules, such as memory allocators, that use them. In this paper, we develop concurrent programming techniques and synchronization primitives, in support of a dynamic memory allocator, that are efficient for use with very high levels of concurrency. We formulate resource allocation as a two-stage process, that decouples accounting for the number of available resources from the tracking of the available resources themselves. To facilitate the accounting stage, we introduce a novel bulk semaphore abstraction that extends traditional semaphore semantics by optimizing for the case where threads operate on the semaphore simultaneously. We also similarly design new collective synchronization primitives that enable groups of cooperating threads to enter critical sections together. Finally, we show that delegation of deferred reclamation to threads already blocked greatly improves efficiency. Using all these techniques, our throughput-oriented memory allocator delivers both high allocation rates and low memory fragmentation on modern GPUs. Our experiments demonstrate that it achieves allocation rates that are on average 16.56 times higher than the counterpart implementation in the CUDA 9 toolkit.