Optimizing threaded MPI execution on SMP clusters

Optimizing threaded MPI execution on SMP clusters
复制标题

DOI:
10.1145/377792.377895
复制
发表时间:
2001-06
影响因子:
0.5
通讯作者:
Hong Tang;Tao Yang
Hong Tang;Tao Yang
中科院分区:
计算机科学4区
文献类型:
--
作者:
Hong Tang;Tao Yang

文献摘要

被引文献

相似文献

我们以前的工作表明,在多程序共享内存机器上使用线程执行MPI程序可以获得很大的性能提升。本文研究了在SMP集群上基于线程的MPI系统的设计与实现。我们的研究表明,与集群环境中基于进程的MPI实现相比,通过适当的线程MPI执行设计,点对点和集体通信性能都可以得到显著改善。我们的贡献包括用于线程MPI执行的层次结构感知和自适应通信方案,以及使用事件驱动同步并提供分离的集体和点对点通信通道的线程安全网络设备抽象。本文描述了我们的设计的实现,并举例说明了它在Linux SMP集群上的性能优势。
Our previous work has shown that using threads to execute MPI programs can yield great performance gain on multiprogrammed shared-memory machines. This paper investigates the design and implementation of a thread-based MPI system on SMP clusters. Our study indicates that with a proper design for threaded MPI execution, both point-to-point and collective communication performance can be improved substantially, compared to a process-based MPI implementation in a cluster environment. Our contribution includes a hierarchy-aware and adaptive communication scheme for threaded MPI execution and a thread-safe network device abstraction that uses event-driven synchronization and provides separated collective and point-to-point communication channels. This paper describes the implementation of our design and illustrates its performance advantage on a Linux SMP cluster.