Compiler-assisted dynamic scheduling for effective parallelization of loop nests on multicore processors

Compiler-assisted dynamic scheduling for effective parallelization of loop nests on multicore processors
复制标题

用于多核处理器上循环嵌套有效并行化的编译器辅助动态调度

DOI:
--
复制
发表时间:
2009
期刊:
ACM SIGPLAN Symposium on Principles & Practice of Parallel Programming
影响因子:
--
通讯作者:
P. Sadayappan
P. Sadayappan
中科院分区:
--
文献类型:
--
作者:
M. Baskaran;N. Vydyanathan;Uday Bondhugula;R. Ramanujam;A. Rountev;P. Sadayappan

文献摘要

被引文献

相似文献

多面体编译技术的最新进展使得自动转换仿射顺序循环嵌套以在多核处理器上平铺并行执行成为可能。然而,对于具有不同维度语句的多语句输入程序,例如Cholesky或LU分解,现有自动并行化方法生成的并行平铺代码可能会遭受严重的负载不平衡,导致多核系统上的可扩展性较差。在本文中,我们开发了一种完全自动的并行化方法,用于将输入仿射顺序代码转换为可以以负载平衡方式在多核系统上执行的高效并行代码。在我们的方法中,我们采用了一种编译时技术,可以在运行时动态提取图块间的依赖关系,并动态调度处理器内核上的并行图块,以改进可扩展的执行。我们的方法消除了程序员干预和重写现有算法的需要,以便在多核上高效并行执行。我们通过使用线性代数计算(LU 和 Cholesky 分解)的比较来证明我们方法的实用性。
Recent advances in polyhedral compilation technology have made it feasible to automatically transform affine sequential loop nests for tiled parallel execution on multi-core processors. However, for multi-statement input programs with statements of different dimensionalities, such as Cholesky or LU decomposition, the parallel tiled code generated by existing automatic parallelization approaches may suffer from significant load imbalance, resulting in poor scalability on multi-core systems. In this paper, we develop a completely automatic parallelization approach for transforming input affine sequential codes into efficient parallel codes that can be executed on a multi-core system in a load-balanced manner. In our approach, we employ a compile-time technique that enables dynamic extraction of inter-tile dependences at run-time, and dynamic scheduling of the parallel tiles on the processor cores for improved scalable execution. Our approach obviates the need for programmer intervention and re-writing of existing algorithms for efficient parallel execution on multi-cores. We demonstrate the usefulness of our approach through comparisons using linear algebra computations: LU and Cholesky decomposition.