An Empirical Study on Deoptimization in the Graal Compiler

An Empirical Study on Deoptimization in the Graal Compiler
复制标题

Graal 编译器中去优化的实证研究

DOI:
10.4230/lipics.ecoop.2017.30
复制
发表时间:
2017
影响因子:
--
通讯作者:
Walter Binder
Walter Binder
中科院分区:
--
文献类型:
--
作者:
Y. Zheng;L. Bulej;Walter Binder

文献摘要

被引文献

相似文献

托管语言平台(如Java虚拟机或公共语言运行时)依赖于动态编译器来实现高性能。除了根据实际的程序执行和底层硬件平台做出优化决策外,动态编译器还处于执行推测性优化的理想位置。然而,这些往往会增加编译成本,因为不成功的推测会触发程序中受影响部分的反优化和重新编译,浪费了之前的工作。尽管推测性优化被广泛使用,但这些优化在额外编译工作方面的成本以前没有被研究过。本文分析了集成在Oracle的HotSpot虚拟机中的Graal动态编译器的行为。我们关注导致程序执行从机器码切换到解释器的情况,并使用三种不同的反优化策略来比较应用程序性能,这三种策略会影响Graal所做的额外编译工作的数量。使用自适应反优化策略,我们设法提高了来自DaCapo、scalabbench和Octane基准套件的基准测试的平均启动性能,主要是通过避免浪费编译工作。在单核系统上,我们观察到DaCapo和scalabbench工作负载的平均加速速度为6.4%,Octane工作负载的平均加速速度为5.1%;随着可用CPU内核数量的增加,性能的提高也随之降低。我们还发现,反优化策略的选择对稳态性能的影响可以忽略不计。这表明投机的代价主要在启动期间产生影响,它会破坏执行程序和编译器之间的微妙平衡,但在稳定状态下很快就会平摊。
Managed language platforms such as the Java Virtual Machine or the Common Language Runtime rely on a dynamic compiler to achieve high performance. Besides making optimization decisions based on the actual program execution and the underlying hardware platform, a dynamic compiler is also in an ideal position to perform speculative optimizations. However, these tend to increase the compilation costs, because unsuccessful speculations trigger deoptimization and recompilation of the affected parts of the program, wasting previous work. Even though speculative optimizations are widely used, the costs of these optimizations in terms of extra compilation work has not been previously studied. In this paper, we analyze the behavior of the Graal dynamic compiler integrated in Oracle's HotSpot Virtual Machine. We focus on situations which cause program execution to switch from machine code to the interpreter, and compare application performance using three different deoptimization strategies which influence the amount of extra compilation work done by Graal. Using an adaptive deoptimization strategy, we managed to improve the average start-up performance of benchmarks from the DaCapo, ScalaBench, and Octane benchmark suites, mostly by avoiding wasted compilation work. On a single-core system, we observed an average speed-up of 6.4% for the DaCapo and ScalaBench workloads, and a speed-up of 5.1% for the Octane workloads; the improvement decreases with an increasing number of available CPU cores. We also find that the choice of a deoptimization strategy has negligible impact on steady-state performance. This indicates that the cost of speculation matters mainly during start-up, where it can disturb the delicate balance between executing the program and the compiler, but is quickly amortized in steady state.