Coarse-Grained Thread Pipelining: A Speculative Parallel Execution Model for Shared-Memory Multiprocessors

Coarse-Grained Thread Pipelining: A Speculative Parallel Execution Model for Shared-Memory Multiprocessors
复制标题

粗粒度线程流水线:共享内存多处理器的推测并行执行模型

DOI:
10.1109/71.954629
复制
发表时间:
2001
期刊:
IEEE Trans. Parallel Distributed Syst.
影响因子:
--
通讯作者:
D. Lilja
D. Lilja
中科院分区:
--
文献类型:
--
作者:
I. Kazi;D. Lilja

文献摘要

被引文献

相似文献

本文提出了一种新的并行化模型,称为粗粒度线程流水线,利用投机粗粒度并行共享内存多处理器系统中的通用应用程序。这种并行化模型,这是基于细粒度的线程流水线模型提出的超线程架构,允许并行执行的循环迭代的流水线方式与运行时数据依赖检查和控制投机。结合运行时依赖性检查的推测性执行允许对不能用现有运行时并行化算法并行化的各种程序构造进行并行化。在这种新技术中,循环迭代的流水线执行导致比其他现有技术更低的并行化开销。我们使用一些真实的应用程序和一个综合基准测试来评估这个新模型的性能。这些实验表明,与并行化开销相比,具有足够大粒度的程序使用此模型获得显着的加速。从合成基准的结果提供了一种手段,用于估计的性能,可以从应用程序,将与此模型并行化。为这个线程流水线模型开发的库例程对于评估由超线程编译器生成的代码的正确性以及调试和验证超线程处理器的模拟器也是有用的。
This paper presents a new parallelization model, called coarse-grained thread pipelining, for exploiting speculative coarse-grained parallelism from general-purpose application programs in shared-memory multiprocessor systems. This parallelization model, which is based on the fine-grained thread pipelining model proposed for the superthreaded architecture, allows concurrent execution of loop iterations in a pipelined fashion with runtime data-dependence checking and control speculation. The speculative execution combined with the runtime dependence checking allows the parallelization of a variety of program constructs that cannot be parallelized with existing runtime parallelization algorithms. The pipelined execution of loop iterations in this new technique results in lower parallelization overhead than in other existing techniques. We evaluated the performance of this new model using some real applications and a synthetic benchmark. These experiments show that programs with a sufficiently large grain size compared to the parallelization overhead obtain significant speedup using this model. The results from the synthetic benchmark provide a means for estimating the performance that can be obtained from application programs that will be parallelized with this model. The library routines developed for this thread pipelining model are also useful for evaluating the correctness of the codes generated by the superthreaded compiler and in debugging and verifying the simulator for the superthreaded processor.