Limits of region-based dynamic binary parallelization

Limits of region-based dynamic binary parallelization
复制标题

基于区域的动态二进制并行化的局限性

DOI:
10.1145/2451512.2451518
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
Edler Von Koch T
Edler Von Koch T
中科院分区:
--
文献类型:
--
作者:
Edler Von Koch T

文献摘要

参考文献

被引文献

相似文献

在由许多小核组成的片上多处理器(CMP)上有效地执行顺序遗留二进制文件是当今最紧迫的问题之一。由于CMP的单核性能较低,单线程执行是一个次优选项,而多线程执行依赖于先前的并行化,这受到针对单核目标编译和优化的应用程序的低级二进制表示的严重阻碍。解决这个问题的最新技术是动态二进制并行化(DBP),它创建了一个虚拟执行环境(VEE),利用底层多核主机来透明地并行化顺序二进制可执行文件。虽然还处于起步阶段,DBP已经在研究界引起了广泛的兴趣。DBP和线程级推测(TLS)的组合使用已经被提出作为一种技术来加速现代CMP上的遗留单处理器代码。在本文中,我们调查DBP的限制,并试图了解这些限制的因素及其实施的成本和管理费用。我们已经进行了广泛的评估,使用可参数化的DBP系统,目标是一个轻量级的架构TLS支持CMP。我们证明,有空间的一个显着减少高达54%的指令的数量上的关键路径的遗留SPEC CPU2006基准。然而,我们发现,这是很难将这些节省转化为实际的性能改进,与现实的硬件支持的实现实现平均1.09的加速比。
Efficiently executing sequential legacy binaries on chip multi-processors (CMPs) composed of many, small cores is one of today's most pressing problems. Single-threaded execution is a suboptimal option due to CMPs' lower single-core performance, while multi-threaded execution relies on prior parallelization, which is severely hampered by the low-level binary representation of applications compiled and optimized for a single-core target. A recent technology to address this problem is Dynamic Binary Parallelization (DBP), which creates a Virtual Execution Environment (VEE) taking advantage of the underlying multicore host to transparently parallelize the sequential binary executable. While still in its infancy, DBP has received broad interest within the research community. The combined use of DBP and thread-level speculation (TLS) has been proposed as a technique to accelerate legacy uniprocessor code on modern CMPs. In this paper, we investigate the limits of DBP and seek to gain an understanding of the factors contributing to these limits and the costs and overheads of its implementation. We have performed an extensive evaluation using a parameterizable DBP system targeting a CMP with light-weight architectural TLS support. We demonstrate that there is room for a significant reduction of up to 54% in the number of instructions on the critical paths of legacy SPEC CPU2006 benchmarks. However, we show that it is much harder to translate these savings into actual performance improvements, with a realistic hardware-supported implementation achieving a speedup of 1.09 on average.
DOI: 10.1109/iiswc.2010.5649169
发表时间: 2010
期刊: --
影响因子: --
作者:
Ioannou N
通讯作者: Ioannou N
单线程应用的芯片多处理器可扩展性
DOI: 10.1145/1105734.1105741
发表时间: 2005
期刊: SIGARCH Comput. Archit. News
影响因子: --
作者:
Neil Vachharajani;M. Iyer;C. Ashok;Manish Vachharajani;David I. August;D. Connors
通讯作者: D. Connors
基于跟踪的 Java JIT 编译器由基于方法的编译器改造而来
DOI: 10.1109/cgo.2011.5764692
发表时间: 2011
期刊: International Symposium on Code Generation and Optimization (CGO 2011)
影响因子: --
作者:
H. Inoue;H. Hayashizaki;Peng Wu;T. Nakatani
通讯作者: T. Nakatani
Java 即时编译器的基于区域的编译技术
DOI: 10.1145/781131.781166
发表时间: 2003
期刊: Proceedings of the ACM SIGPLAN 2003 conference on Programming language design and implementation
影响因子: --
作者:
T. Suganuma;T. Yasue;T. Nakatani
通讯作者: T. Nakatani
在动态二进制翻译器中使用并行任务场的广义即时跟踪编译
DOI: 10.1145/1993498.1993508
发表时间: 2011
期刊: --
影响因子: --
作者:
Böhm I
通讯作者: Böhm I