Tight analysis of the performance potential of thread speculation using spec CPU 2006

Tight analysis of the performance potential of thread speculation using spec CPU 2006
复制标题

使用规格 CPU 2006 严格分析线程推测的性能潜力

DOI:
--
复制
发表时间:
2007
期刊:
ACM SIGPLAN Symposium on Principles & Practice of Parallel Programming
影响因子:
--
通讯作者:
C. Polychronopoulos
C. Polychronopoulos
中科院分区:
--
文献类型:
--
作者:
A. Kejariwal;Xinmin Tian;M. Girkar;Wei Li;Sergey Kozhukhov;U. Banerjee;A. Nicolau;A. Veidenbaum;C. Polychronopoulos

文献摘要

被引文献

相似文献

诸如Intel®1 Core™2 Duo处理器的多核促进普通程序的高效线程级并行执行,其中不同的执行线程被映射到不同的物理处理器上。在这种情况下,已经提出了几种技术的程序的自动并行化。最近,线程级推测(TLS)已被提出作为一种手段来并行化难以分析的串行代码。通常,可以采用一种以上的技术来并行化给定的程序。各种技术的适用性的重叠性质使得很难评估每种技术的内在性能潜力。在本文中,我们提出了一个严格的分析(独特的)性能潜力:(a)TLS的一般和(B)特定类型的线程级投机,即,控制推测,数据依赖推测和数据值推测,为SPEC 2 CPU 2006基准测试套件,根据各种限制因素,如线程开销和误推测惩罚。据我们所知,这是基于SPEC CPU 2006的TLS的第一次评估,并考虑了上述现实生活中的约束。我们的分析表明,在最内层的循环级别,上界的加速唯一可实现的通过TLS与国家的最先进的线程实现的SPEC CINT 2006和CFP 2006是1%的顺序。
Multi-cores such as the Intel®1 Core™2 Duo processor, facilitate efficient thread-level parallel execution of ordinary programs, wherein the different threads-of-execution are mapped onto different physical processors. In this context, several techniques have been proposed for auto-parallelization of programs. Recently, thread-level speculation (TLS) has been proposed as a means to parallelize difficult-to-analyze serial codes. In general, more than one technique can be employed for parallelizing a given program. The overlapping nature of the applicability of the various techniques makes it hard to assess the intrinsic performance potential of each. In this paper, we present a tight analysis of the (unique) performance potential of both: (a) TLS in general and (b) specific types of thread-level speculation, viz., control speculation, data dependence speculation and data value speculation, for the SPEC2 CPU2006 benchmark suite in light of the various limiting factors such as the threading overhead and misspeculation penalty. To the best of our knowledge, this is the first evaluation of TLS based on SPEC CPU2006 and accounts for the aforementioned real-life con-straints. Our analysis shows that, at the innermost loop level, the upper bound on the speedup uniquely achievable via TLS with the state-of-the-art thread implementations for both SPEC CINT2006 and CFP2006 is of the order of 1%.