Automatic Handling of Global Variables for Multi-threaded MPI Programs

Automatic Handling of Global Variables for Multi-threaded MPI Programs
复制标题

多线程 MPI 程序的全局变量的自动处理

DOI:
10.1109/icpads.2011.33
复制
发表时间:
2011
期刊:
2011 IEEE 17th International Conference on Parallel and Distributed Systems
影响因子:
--
通讯作者:
E. Rodrigues
E. Rodrigues
中科院分区:
--
文献类型:
--
作者:
G. Zheng;Stas Negara;C. Mendes;L. Kalé;E. Rodrigues

文献摘要

被引文献

相似文献

MPI 标准的传统实现倾向于将每个处理器关联一个 MPI 进程,这限制了它们对现代多核平台的支持。一种越来越流行的方法是将 MPI 与线程结合起来,其中 MPI“进程”是轻量级线程。然而,传统 MPI 应用程序中的全局变量提出了一个挑战,因为它们可能被多个 MPI 线程同时访问。因此,在此类 MPI 执行环境中将遗留 MPI 应用程序转变为线程安全的需要正确处理全局变量。在本文中,我们提出了三种自动消除全局变量的方法,以确保 MPI 程序的线程安全。这些方法包括:(a)基于编译器的重构技术,以基于 Photran 的工具为例,它可以自动执行 Fortran 编写的程序的源到源转换,(b)基于全局偏移表(GOT)的技术; (c)基于线程本地存储(TLS)的技术。第二种和第三种方法自动检测全局变量并在运行时为每个线程将它们私有化。我们讨论这些方法的优点和缺点,并使用综合基准(例如 NAS 基准)和真实的科学应用程序(FLASH 代码)比较它们的性能。
Conventional implementations of the MPI standard tend to associate one MPI process per processor, which limits their support for modern multi-core platforms. An increasingly popular approach is to combine MPI with threads where MPI "processes" are light-weight threads. Global variables in legacy MPI applications, however, present a challenge because they may be accessed by multiple MPI threads simultaneously. Thus, transforming legacy MPI applications to become thread-safe in such MPI execution environments requires proper handling of global variables. In this paper, we present three approaches to automatically eliminate global variables to ensure thread-safety for an MPI program. These approaches include: (a) a compiler-based refactoring technique, using a Photran-based tool as an example, which automates the source-to-source transformation for programs written in Fortran, (b) a technique based on a global offset table (GOT); and (c) a technique based on thread local storage (TLS). The second and third methods automatically detect global variables and privatize them for each thread at runtime. We discuss the advantages and disadvantages of these approaches and compare their performance using both synthetic benchmarks, such as the NAS Benchmarks, and a real scientific application, the FLASH code.