An automatic object inlining optimization and its evaluation

An automatic object inlining optimization and its evaluation
复制标题

一种自动对象内联优化及其评估

DOI:
10.1145/349299.349344
复制
发表时间:
2000
期刊:
Proceedings of the 2013 International Conference on Principles and Practices of Programming on the Java Platform: Virtual Machines, Languages, and Tools
影响因子:
--
通讯作者:
A. Chien
A. Chien
中科院分区:
--
文献类型:
--
作者:
Julian T Dolby;A. Chien

文献摘要

被引文献

相似文献

自动对象内联[19,20]通过将父对象和子对象融合在一起来转换堆数据结构。它可以通过减少对象分配和指针解引用开销来提高运行时间。我们报告继续研究对象内联优化的工作。特别是,我们提出了一个新的语义推导对象内联的正确性条件,和程序分析,扩展了我们以前的工作。提出了一种对象内联转换算法,重点讨论了一种优化类字段布局以减少代码扩展的新算法。最后,我们详细介绍了11个程序和库(包括Xpdf,25,000行便携式文档格式(PDF)文件浏览器),利用硬件措施的影响,内存系统的一个更全面的评估。 我们表明,我们的分析规模有效地大型程序,发现许多inlinable领域(45在xpdf)在可接受的成本,和weeshow,在某些程序中,它发现几乎所有的领域,对象内联是正确的,平均40%的这些领域在我们的基准。我们实现了我们的分析在一个先进的分析基础设施,我们表明,相比传统的1-CFA,该基础设施提供了更好的结果和更低,更具可扩展性的成本。在所有的程序中,分析发现平均大约30%的对象是可内联的。我们的转换只增加了20%的代码大小,同时内联了这30%的字段。内联这些对象平均消除了28%的字段读取,58%的对象创建,12%的所有加载。此外,优化的程序显著改善了内存引用行为,使L1数据缓存未命中减少25%,读取暂停减少25%。平均运行时间提高了14%。
Automatic object inlining [19, 20] transforms heapdata structures by fusing parent and child objects together. It canimprove runtime by reducing object allocation and pointer dereferencecosts. We report continuing work studying object inlining optimizations. In particular, we present a new semantic derivation of the correctness conditions for object inlining, and program analysis which extends our previous work. And we present an object inlining transformation, focusing on a new algorithm which optimizes class field layout to minimize code expansion. Finally, we detail a fuller evaluation on eleven programs and libraries (including Xpdf, the 25,000 line Portable Document Format (PDF) file browser) that utilizes hardware measures of impact on the memory system. We show that our analysis scales effectively to large programs,finding many inlinable fields (45 in xpdf) at acceptable cost, and weshow that, on some programs, it finds nearly all fields for which object inlining is correct, and averages 40% of such fields across our benchmarks. We implement our analyses in an advanced analysis infrastructure, and we show that, compared to traditional 1-CFA, that infrastructure provides better results and lower and more scalable cost. Across all programs, analysis identified about 30% of objects as inlinable on average. Our transformation increases code size by only 20% while inlining this 30% of fields. Inlining these objects eliminated on average 28% of field reads, 58% of object creations, 12% of all loads. Further, the optimized programs have significantly improved memory reference behavior, producing 25% fewer L1 data cache misses and 25% fewer read stalls. On average the runtime improved by 14%.