Exploiting hierarchy in parallel computer networks to optimize collective operation performance

Exploiting hierarchy in parallel computer networks to optimize collective operation performance
复制标题

DOI:
10.1109/ipdps.2000.846009
复制
发表时间:
2000-02
期刊:
Proceedings 14th International Parallel and Distributed Processing Symposium. IPDPS 2000
影响因子:
--
通讯作者:
N. Karonis;B. Supinski;Ian T Foster;W. Gropp;E. Lusk;J. Bresnahan
N. Karonis;B. Supinski;Ian T Foster;W. Gropp;E. Lusk;J. Bresnahan
中科院分区:
其他
文献类型:
--
作者:
N. Karonis;B. Supinski;Ian T Foster;W. Gropp;E. Lusk;J. Bresnahan

文献摘要

被引文献

相似文献

集体传播行动的有效实施受到了广泛关注。最初的努力建立了网络通信模型,并基于这些模型生成了“最优”树。然而,这些初始工作所使用的模型假定任意两个进程之间的点对点延迟是相等的。这一假设在异构系统(如smp集群和广域“计算网格”)中是违反的,因此,利用这些模型生成的树的集体操作执行得不是最优。作为回应,最近的工作集中在为集体操作创建拓扑感知树,以最大限度地减少跨较慢通道(例如广域网)的通信。虽然这些努力具有显著的通信优势,但它们都将网络的视图限制在两层。我们提出了一种基于网络多层视图的策略。通过创建多层拓扑树,我们利用了网络中每一层的通信成本差异。我们使用该策略在MPICH- g中实现了几个MPI集合操作的拓扑感知版本,MPICH- g是流行的MPI标准的MPICH实现的支持globus的版本。使用Globus发现的拓扑信息,我们在执行过程中自动构建这些拓扑感知树,从而使MPI应用程序程序员不必编写特殊的文件或函数来向MPICH库描述拓扑。我们通过将多层方法与MPICH提供的默认(拓扑不知情)实现和拓扑感知的两层实现进行比较,展示了多层方法的优势。
The efficient implementation of collective communication operations has received much attention. Initial efforts modeled network communication and produced "optimal" trees based on those models. However, the models used by these initial efforts assumed equal point-to-point latencies between any two processes. This assumption is violated in heterogeneous systems such as clusters of SMPs and wide-area "computational grids", and as a result, collective operations that utilize the trees generated by these models perform suboptimally. In response, more recent work has focused on creating topology-aware trees for collective operations that minimize communication across slower channels (e.g., a wide-area network). While these efforts have significant communication benefits, they all limit their view of the network to only two layers. We present a strategy based upon a multilayer view of the network. By creating multilevel topology trees we take advantage of communication cost differences at every level in the network. We used this strategy to implement topology-aware versions of several MPI collective operations in MPICH-G, the Globus-enabled version of the popular MPICH implementation of the MPI standard. Using information about topology discovered by Globus, we construct these topology-aware trees automatically during execution, thus freeing the MPI application programmer from having to write special files or functions to describe the topology to the MPICH library. We present results demonstrating the advantages of our multilevel approach by comparing it to the default (topology-unaware) implementation provided by MPICH and a topology-aware two-layer implementation.