The study and handling of program inputs in the selection of garbage collectors

The study and handling of program inputs in the selection of garbage collectors
复制标题

垃圾收集器选择中程序输入的研究和处理

DOI:
--
复制
发表时间:
2009
期刊:
OPSR
影响因子:
--
通讯作者:
E. Zhang
E. Zhang
中科院分区:
--
文献类型:
--
作者:
Xipeng Shen;Feng Mao;Kai Tian;E. Zhang

文献摘要

被引文献

相似文献

许多研究表明,对于不同的应用程序,一组垃圾收集器中表现最好的往往是不同的。研究人员提出了特定于应用程序的垃圾收集器选择。在这项工作中,我们集中在第二个方面的问题:垃圾收集器的选择程序输入的影响。我们为一组Java基准测试收集了数十到数百个输入,并在具有不同堆大小和垃圾收集器的Jikes RVM上测量了它们的性能。严格的统计分析产生了四重见解。首先,输入会显著影响垃圾收集器的相对性能,从而导致垃圾收集器的顶级集合在输入之间存在很大差异。因此,分析一次或几次运行不足以选择适合大多数输入的垃圾收集器。第二,当堆大小比率固定时,一两种类型的垃圾收集器足以刺激程序在所有输入上的最佳性能。第三,对于某些程序,堆大小比会显著影响不同类型垃圾收集器的相对性能。为了在这些程序上选择垃圾收集器,必须有一个交叉输入预测模型,该模型可以预测在任意输入上执行的最小可能堆大小。最后,通过统计学习技术,我们研究了交叉输入影响的可预测性。实验结果表明,使用回归和分类技术,在给定应用程序的任意输入的情况下,可以以合理的准确度预测最佳垃圾收集器(沿着最小可能的堆大小)。这种探索为定制垃圾收集器的选择提供了机会,不仅适用于应用程序,而且适用于它们的输入。
Many studies have shown that the best performer among a set of garbage collectors tends to be different for different applications. Researchers have proposed applicationspecific selection of garbage collectors. In this work, we concentrate on a second dimension of the problem: the influence of program inputs on the selection of garbage collectors. We collect tens to hundreds of inputs for a set of Java benchmarks, and measure their performance on Jikes RVM with different heap sizes and garbage collectors. A rigorous statistical analysis produces four-fold insights. First, inputs influence the relative performance of garbage collectors significantly, causing large variations to the top set of garbage collectors across inputs. Profiling one or few runs is thus inadequate for selecting the garbage collector that works well for most inputs. Second, when the heap size ratio is fixed, one or two types of garbage collectors are enough to stimulate the top performance of the program on all inputs. Third, for some programs, the heap size ratio significantly affects the relative performance of different types of garbage collectors. For the selection of garbage collectors on those programs, it is necessary to have a cross-input predictive model that predicts the minimum possible heap size of the execution on an arbitrary input. Finally, by adoptingstatistical learning techniques, we investigate the cross-input predictability of the influence. Experimental results demonstrate that with regression and classification techniques, it is possible to predict the best garbage collector (along with the minimum possible heap size) with reasonable accuracy given an arbitrary input to an application. The exploration opens the opportunities for tailoring the selection of garbage collectors to not only applications but also their inputs.