Massively parallel GPU memory compaction

Massively parallel GPU memory compaction
复制标题

大规模并行 GPU 内存压缩

DOI:
10.1145/3315573.3329979
复制
发表时间:
2019
期刊:
Proceedings of the 2019 ACM SIGPLAN International Symposium on Memory Management
影响因子:
--
通讯作者:
Hidehiko Masuhara
Hidehiko Masuhara
中科院分区:
--
文献类型:
--
作者:
Matthias Springer;Hidehiko Masuhara

文献摘要

相似文献

内存碎片是动态内存分配器中一个被广泛研究的问题。众所周知,碎片可能导致过早的内存不足错误和较差的缓存性能。随着SIMD加速器动态内存分配器的出现,内存碎片成为这类体系结构上日益重要的问题。然而,到目前为止,它几乎没有受到关注。基于SIMD架构(如gpu)的内存绑定应用程序可能会由于效率较低的矢量加载/存储指令而经历额外的减速。我们提出了CompactGpu,一个增量,全并行,就地内存碎片整理系统的gpu。CompactGpu是dynasar动态内存分配器的扩展,通过合并部分占用的内存块,以完全并行的方式对堆进行碎片整理。我们开发了几种在SIMD/GPU架构上有效的内存碎片整理实现技术,例如查找碎片整理块候选项和基于位图的快速指针重写。基准测试表明,我们的实现非常快,通常具有比压缩开销更高的性能增益。它还可以减少总体内存使用。
Memory fragmentation is a widely studied problem of dynamic memory allocators. It is well known that fragmentation can lead to premature out-of-memory errors and poor cache performance.With the recent emergence of dynamic memory allocators for SIMD accelerators, memory fragmentation is becoming an increasingly important problem on such architectures. Nevertheless, it has received little attention so far. Memory-bound applications on SIMD architectures such as GPUs can experience an additional slowdown due to less efficient vector load/store instructions.We propose CompactGpu, an incremental, fully-parallel, in-place memory defragmentation system for GPUs. CompactGpu is an extension to the DynaSOAr dynamic memory allocator and defragments the heap in a fully parallel fashion by merging partly occupied memory blocks. We developed several implementation techniques for memory defragmentation that are efficient on SIMD/GPU architectures, such as finding defragmentation block candidates and fast pointer rewriting based on bitmaps.Benchmarks indicate that our implementation is very fast with typically higher performance gains than compaction overheads. It can also decrease the overall memory usage.