Runtime automatic speculative parallelization

Runtime automatic speculative parallelization
复制标题

运行时自动推测并行化

DOI:
10.1109/cgo.2011.5764675
复制
发表时间:
2011
期刊:
International Symposium on Code Generation and Optimization (CGO 2011)
影响因子:
--
通讯作者:
K. Olukotun
K. Olukotun
中科院分区:
--
文献类型:
--
作者:
Ben Hertzberg;K. Olukotun

文献摘要

被引文献

相似文献

我们提出了一种自动推测性线程化(RASP),一种以用户透明的方式从运行的应用程序中动态提取推测性线程的技术。通过利用CMP中的空闲内核来分析、优化和参与正在运行的顺序程序的执行,RASP使一组更简单的内核能够实现与更复杂的内核相当的顺序性能。与其他自动推测并行化系统相比,RASP使用动态二进制翻译来动态优化应用程序,而无需重新编译或源代码。RASP实现这些加速而不依赖于特殊用途的硬件支持; RASP的动态分析使用了一个聪明的变化对传统的性能监控,而RASP的投机执行依赖于相同的简单的硬件支持的投机,已提出简化并行编程。在一个由四个有序内核组成的模拟集群上,RASP将SPEC 2006整数基准测试平均加速了49%,对于科学和多媒体工作负载也有很好的效果。
We present Runtime Automatic Speculative Parallelization (RASP), a technique for the dynamic extraction of speculative threads from a running application in a user-transparent fashion. By leveraging the idle cores in a CMP to analyze, optimize, and participate in the execution of a running sequential program, RASP enables a collection of simpler cores to achieve sequential performance on par with a significantly more complex core. In contrast to other systems for automatic speculative parallelization, RASP uses dynamic binary translation to optimize applications on-the-fly, without any need for recompilation or source code. RASP achieves these speedups without relying on special-purpose hardware support; RASP's dynamic profiling uses a clever variation on conventional performance monitoring, while RASP's speculative execution relies on the same simple hardware support for speculation that has been proposed for simplifying parallel programming. On a simulated cluster of four in-order cores, RASP accelerates SPEC2006 integer benchmarks by an average of 49%, with promising results for scientific and multimedia workloads as well.