BOLT: A Practical Binary Optimizer for Data Centers and Beyond

BOLT: A Practical Binary Optimizer for Data Centers and Beyond
复制标题

BOLT:适用于数据中心及其他领域的实用二进制优化器

DOI:
10.1109/cgo.2019.8661201
复制
发表时间:
2018
期刊:
2019 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
影响因子:
--
通讯作者:
Guilherme Ottoni
Guilherme Ottoni
中科院分区:
--
文献类型:
--
作者:
Maksim Panchenko;R. Auler;B. Nell;Guilherme Ottoni

文献摘要

被引文献

相似文献

随着计算继续向数据中心转移,对大规模应用的性能优化最近变得越来越重要。数据中心的应用程序通常非常大且复杂,这使得代码布局成为提高其性能的重要优化。这激发了最近对实用技术的调查,以改善编译时间和链接时间的代码布局。尽管后链路优化器过去在过去的成功中取得了一些成功,但在现代数据中心应用程序的背景下,最近没有探索其收益。在本文中,我们提出了Bolt,这是在LLVM框架之上构建的开源后链接优化器。使用基于样本的分析,Bolt即使是使用反馈驱动的优化(FDO)和链接时间优化(LTO)构建的高度优化二进制文件,也可以提高现实世界应用的性能。我们证明,即使在整个程序级别和概要信息的存在下完成后者,后链路的性能改进也与常规编译器的优化互补。我们评估了Facebook数据中心工作负载和开源编译器的螺​​栓。对于数据中心的应用程序,Bolt在配置文件引导的重新排序和LTO之上达到了高达7.0%的性能加速。对于GCC和CLANG编译器,我们的评估表明,螺栓在FDO和LTO的顶部将其二进制速度提高高达20.4%,如果二进制文件是没有FDO和LTO的,则螺栓的二进制速度高达52.1%。
Performance optimization for large-scale applications has recently become more important as computation continues to move towards data centers. Data-center applications are generally very large and complex, which makes code layout an important optimization to improve their performance. This has motivated recent investigation of practical techniques to improve code layout at both compile time and link time. Although post-link optimizers had some success in the past, no recent work has explored their benefits in the context of modern data-center applications. In this paper, we present BOLT, an open-source post-link optimizer built on top of the LLVM framework. Utilizing sample-based profiling, BOLT boosts the performance of real-world applications even for highly optimized binaries built with both feedback-driven optimizations (FDO) and link-time optimizations (LTO). We demonstrate that post-link performance improvements are complementary to conventional compiler optimizations, even when the latter are done at a whole-program level and in the presence of profile information. We evaluated BOLT on both Facebook data-center workloads and open-source compilers. For data-center applications, BOLT achieves up to 7.0% performance speedups on top of profile-guided function reordering and LTO. For the GCC and Clang compilers, our evaluation shows that BOLT speeds up their binaries by up to 20.4% on top of FDO and LTO, and up to 52.1% if the binaries are built without FDO and LTO.