Extended Memory Reuse: An Optimisation for Reducing Memory Allocations

Extended Memory Reuse: An Optimisation for Reducing Memory Allocations
复制标题

扩展内存重用:减少内存分配的优化

DOI:
--
复制
发表时间:
2018
期刊:
International Symposium on Implementation and Application of Functional Languages
影响因子:
--
通讯作者:
S. Scholz
S. Scholz
中科院分区:
--
文献类型:
--
作者:
Hans;Artjoms Šinkarovs;S. Scholz

文献摘要

参考文献

被引文献

相似文献

在本文中,我们提出了一种基于引用计数的垃圾收集优化方法。优化旨在减少对堆管理器的调用总数,同时保留引用计数的主要好处,即无需全局垃圾收集即可实现就地更新和内存释放。其关键思想是小心地延长变量的生命周期,以便在相同大小的内存分配之后的内存释放可以被直接的内存重用所取代。事实证明,这种内存重用在计算密集型应用程序的最内部循环的上下文中特别有用。它导致了在缓冲区之间执行指针交换的运行时行为,其方式与在需要显式内存管理的语言中手动实现的方式相同,例如C。我们已经在单任务C编译器工具链的上下文中实现了所提出的优化。本文提供了我们的优化的算法描述,并在一系列基准测试上对其有效性进行了评估,包括Rodinia基准测试的子集和NAS并行基准测试。我们表明,对于循环内具有分配的几个基准测试,我们的优化将分配的数量减少了几个数量级。我们也没有观察到对总体内存占用或总体运行时的负面影响。相反,对于一些顺序执行,我们发现略有改进,在GPU设备上,我们观察到加速比高达4倍。
In this paper we present an optimisation for reference counting based garbage collection. The optimisation aims at reducing the total number of calls to the heap manager while preserving the key benefits of reference counting, i.e. the opportunities for in-place updates as well as memory deallocation without global garbage collection. The key idea is to carefully extend the lifetime of variables so that memory deallocations followed by memory allocations of the same size can be replaced by a direct memory reuse. Such memory reuse turns out particularly useful in the context of innermost loops of compute-intensive applications. It leads to a runtime behaviour that performs pointer swaps between buffers in the same way it would be implemented manually in languages that require explicit memory management, e.g. C. We have implemented the proposed optimisation in the context of the Single-Assignment C compiler tool chain. The paper provides an algorithmic description of our optimisation and an evaluation of its effectiveness over a collection of benchmarks including a subset of the Rodinia benchmarks and the NAS Parallel Benchmarks. We show that for several benchmarks with allocations within loops our optimisation reduces the amount of allocations by a few orders of magnitude. We also observe no negative impact on the overall memory footprint nor on the overall runtime. Instead, for some sequential executions we find mild improvement, and on GPU devices we observe speedups of up to a factor of 4x.
并行表演的抽象表现主义
DOI: 10.1145/2774959.2774962
发表时间: 2015
期刊: --
影响因子: --
作者:
Bernecky R
通讯作者: Bernecky R