Automatic Coarse Grain Task Parallel Processing on SMP Using OpenMP

Automatic Coarse Grain Task Parallel Processing on SMP Using OpenMP
复制标题

使用 OpenMP 在 SMP 上自动进行粗粒度任务并行处理

DOI:
10.1007/3-540-45574-4_13
复制
发表时间:
2000
期刊:
Proceedings Seventh Heterogeneous Computing Workshop (HCW'98)
影响因子:
--
通讯作者:
K. Ishizaka
K. Ishizaka
中科院分区:
--
文献类型:
--
作者:
H. Kasahara;M. Obata;K. Ishizaka

文献摘要

参考文献

被引文献

相似文献

提出了一种简单高效的分级粗粒度任务并行处理方案在SMP机上的实现方法。OSCAR多粒度并行化编译器自动生成包含OpenMP指令的并行化代码,并在商业SMP机器上进行了性能评估。粗粒度任务并行处理对于提高从单片多处理机到高性能计算机的各种多处理机系统的有效性能具有重要意义。该方案将Fortran程序分解为粗粒度任务,通过考虑控制依赖和数据依赖的最早可执行条件分析来分析任务间的并行性,将粗粒度任务静态调度到线程或生成动态任务调度代码将任务分配给线程,并为SMP机器生成OpenMP Fortran源代码。使用OSCAR编译器生成的OpenMP生成的线程并行代码在程序开始时只派生一次线程,在结束时只连接一次线程,即使程序是基于分层粗粒度任务并行处理的概念进行并行处理的。该方案在8处理器SMP机器IBM Rs6000 SP 604e High Node上使用OSCAR多粒度编译器新开发的OpenMP后端进行了性能评估。评估表明,对于SPEC 95fp Swin、TOMCATV、HYDRO2D、MGRID和Perfect基准ARC2D,使用IBM XL Fortran编译器5.1版的OSCAR编译器的加速比是原生XL Fortran编译器的1.5到3倍。
This paper proposes a simple and efficient implementation method for a hierarchical coarse grain task parallel processing scheme on a SMP machine. OSCAR multigrain parallelizing compiler automatically generates parallelized code including OpenMP directives and its performance is evaluated on a commercial SMP machine. The coarse grain task parallel processing is important to improve the effective performance of wide range of multiprocessor systems from a single chip multiprocessor to a high performance computer beyond the limit of the loop parallelism. The proposed scheme decomposes a Fortran program into coarse grain tasks, analyzes parallelism among tasks by "Earliest Executable Condition Analysis" considering control and data dependencies, statically schedules the coarse grain tasks to threads or generates dynamic task scheduling codes to assign the tasks to threads and generates OpenMP Fortran source code for a SMP machine. The thread parallel code using OpenMP generated by OSCAR compiler forks threads only once at the beginning of the program and joins only once at the end even though the program is processed in parallel based on hierarchical coarse grain task parallel processing concept. The performance of the scheme is evaluated on 8-processor SMP machine, IBM RS6000 SP 604e High Node, using a newly developed OpenMP backend of OSCAR multigrain compiler. The evaluation shows that OSCAR compiler with IBM XL Fortran compiler version 5.1 gives us 1.5 to 3 times larger speedup than the native XL Fortran compiler for SPEC 95fp SWIM, TOMCATV, HYDRO2D, MGRID and Perfect Benchmarks ARC2D.
关于决定无法决定的事情
DOI: --
发表时间: 2004
期刊:
影响因子: --
作者:
坂元慶行;中村 隆;前田忠彦;土屋隆裕;Seigo UENO;立岩真也;立岩真也;立岩真也;立岩真也;立岩真也;立岩真也;立岩真也;後藤玲子;Shinya TATEIWA;Shinya TATEIWA;Shinya TATEIWA;Shinya TATEIWA;Shinya TATEIWA;Reiko GOTOH;Seigo UENO;立岩 真也;立岩 真也;立岩 真也;立岩真也;立岩真也;立岩真也;立岩真也;Shinya TATEIWA;Shinya TATEIWA;Shinya TATEIWA;Shinya TATEIWA;Shinya TATEIWA;立岩 真也;立岩 真也;立岩 真也;立岩真也;Shinya TATEIWA;Reiko GOTOH;Seigo UENO;立岩 真也;立岩 真也;立岩 真也;立岩 真也;後藤玲子;立岩真也;Shinya TATEIWA;立岩真也;Shinya TATEIWA;立岩真也;Shinya TATEIWA;立岩真也;立岩 真也;立岩真也;Shinya TATEIWA;後藤玲子;Reiko GOTOH;後藤玲子;Reiko GOTOH;後藤玲子;Reiko GOTOH;後藤玲子;Reiko GOTOH;立岩真也;Shinya TATEIWA;立岩真也;Shinya TATEIWA;樋澤吉彦;Yoshihiko HIZAWA;立岩真也;Shinya TATEIWA;立岩真也;Shinya TATEIWA;立岩真也;Shinya TATEIWA;立岩真也;Shinya TATEIWA
通讯作者: Shinya TATEIWA