Barriers: friend or foe?

Barriers: friend or foe?
复制标题

障碍:朋友还是敌人?

DOI:
10.1145/1029873.1029891
复制
发表时间:
2004
期刊:
Proceedings. 26th International Conference on Software Engineering
影响因子:
--
通讯作者:
Antony Hosking
Antony Hosking
中科院分区:
--
文献类型:
--
作者:
S. Blackburn;Antony Hosking

文献摘要

被引文献

相似文献

现代垃圾收集器依靠突变器对堆访问的读写障碍,跟踪收集堆垃圾的不同区域之间的参考,并将突变器的动作与收集器的动作同步。这是一个长期的未经测试的假设,即障碍对垃圾收集的应用施加了明显的开销。结果,研究人员致力于开发消除不必要的障碍的优化方法,或提出了用于垃圾收集的新算法,以避免需要障碍,同时保留独立收集堆分区的能力。根据此处介绍的结果,我们消除了障碍开销应该是这种努力的主要动机的假设。 我们提出了一种用于精确测量突变器开销的方法,用于与突变器堆访问相关的屏障。我们提供了不同风格的障碍物分类法,并测量Jikes RVM中不同垃圾收集器的一系列流行障碍的成本。我们的结果表明,尽管结果因建筑而异,但障碍对突变器的成本却令人惊讶。我们发现,合理的世代写入障碍的平均开销平均不到2%,在最坏情况下不到6%。此外,我们发现,读取屏障的平均开销仅由PowerPC上读取的低阶位的无条件面膜组成,仅为0.85%,而在AMD上为8.05%。在读写障碍的情况下,我们发现二阶位置效应有时比障碍本身的开销更重要,从而导致许多情况下的违反直觉加速。
Modern garbage collectors rely on read and write barriers imposed on heap accesses by the mutator, to keep track of references between different regions of the garbage collected heap, and to synchronize actions of the mutator with those of the collector. It has been a long-standing untested assumption that barriers impose significant overhead to garbage-collected applications. As a result, researchers have devoted effort to development of optimization approaches for elimination of unnecessary barriers, or proposed new algorithms for garbage collection that avoid the need for barriers while retaining the capability for independent collection of heap partitions. On the basis of the results presented here, we dispel the assumption that barrier overhead should be a primary motivator for such efforts. We present a methodology for precise measurement of mutator overheads for barriers associated with mutator heap accesses. We provide a taxonomy of different styles of barrier and measure the cost of a range of popular barriers used for different garbage collectors within Jikes RVM. Our results demonstrate that barriers impose surprisingly low cost on the mutator, though results vary by architecture. We found that the average overhead for a reasonable generational write barrier was less than 2% on average, and less than 6% in the worst case. Furthermore, we found that the average overhead of a read barrier consisting of just an unconditional mask of the low order bits read on the PowerPC was only 0.85%, while on the AMD it was 8.05%. With both read and write barriers, we found that second order locality effects were sometimes more important than the overhead of the barriers themselves, leading to counter-intuitive speedups in a number of situations.