Evaluation of MPI implementations on grid-connected clusters using an emulated WAN environment

Evaluation of MPI implementations on grid-connected clusters using an emulated WAN environment
复制标题

使用模拟 WAN 环境评估并网集群上的 MPI 实施

DOI:
10.1109/ccgrid.2003.1199347
复制
发表时间:
2003
期刊:
CCGrid 2003. 3rd IEEE/ACM International Symposium on Cluster Computing and the Grid, 2003. Proceedings.
影响因子:
--
通讯作者:
Y. Ishikawa
Y. Ishikawa
中科院分区:
--
文献类型:
--
作者:
Motohiko Matsuda;T. Kudoh;Y. Ishikawa

文献摘要

被引文献

相似文献

MPICHG-2库中集成了用于集群计算的MPICH-SCore高性能通信库,以使PC集群适应网格环境。集成库称为MPICH-G2/SCore。此外,为了与其他方法进行比较,对MPICH-SCore本身进行了扩展,将其网络数据包封装为UDP数据包,以便通过L3交换机传递数据包。这个扩展被称为UDP封装的MPICH-SCore。在本文中,MPI库,UDP封装的MPICH-SCore,MPICH-G2/SCore,和MPICH-P4的三种实现,使用模拟WAN环境,其中两个集群,每个集群由16台主机,由路由器PC连接进行评估。路由器PC控制集群之间消息传递的延迟,并且增加的延迟在往返时间中从1毫秒到4毫秒变化。使用NAS并行基准测试进行实验,结果表明UDP封装的MPICH-SCore通常比其他实现更好。然而,这些差异对基准来说并不重要。初步结果表明,LU基准测试的性能线性扩展,往返延迟低于4毫秒。CG和MG基准测试分别显示了1.13和1.24倍的可扩展性,4毫秒的往返延迟。
The MPICH-SCore high performance communication library for cluster computing is integrated into the MPICHG-2 library in order to adapt PC clusters to a Grid environment. The integrated library is called MPICH-G2/SCore. In addition, for the purpose of comparison with other approaches, MPICH-SCore itself is extended to encapsulate its network packet into a UDP packet so that packets are delivered via L3 switches. This extension is called UDP-encapsulated MPICH-SCore. In this paper, three implementations of the MPI library, UDP-encapsulated MPICH-SCore, MPICH-G2/SCore, and MPICH-P4, are evaluated using an emulated WAN environment where two clusters, each consisting of sixteen hosts, are connected by a router PC. The router PC controls the latency of message delivery between clusters, and the added latency is varied from I millisecond to 4 milliseconds in round-trip time. Experiments are performed using the NAS Parallel Benchmarks, which show UDP-encapsulated MPICH-SCore most often performs better than other implementations. However, the differences are not critical for the benchmarks. The preliminary results show that the performance of the LU benchmark scales up linearly with under 4 millisecond round-trip latency. The CG and MG benchmarks show the scalability of 1.13 and 1.24 times with 4 millisecond round-trip latency, respectively.