Efficient shared memory and RDMA based collectives on multi-rail QsNetII SMP clusters

Efficient shared memory and RDMA based collectives on multi-rail QsNetII SMP clusters
复制标题

DOI:
10.1007/s10586-008-0065-8
复制
发表时间:
2008-12
期刊:
Cluster Computing
影响因子:
--
通讯作者:
Y. Qian;A. Afsahi
Y. Qian;A. Afsahi
中科院分区:
其他
文献类型:
--
作者:
Y. Qian;A. Afsahi

文献摘要

被引文献

相似文献

对称多处理器(SMP)集群在实现高性能方面比以往任何时候都更加普遍。在集群上运行的科学应用程序广泛使用集体通信。多轨网络上的共享存储器通信和远程直接存储器访问(RDMA)是解决对节点内和节点间通信的日益增长的需求的有前途的方法,从而提高新兴的多核SMP集群中的集体的性能。在这方面,本文设计和评估两类集体通信算法直接在Elan用户级的多轨二次QsNetII与消息条带:1)基于RDMA的传统多端口算法,用于收集、全收集和全对全集合,用于中到大消息,2)针对中小规模报文的基于RDMA和SMP感知的多端口全聚集算法。对于4KB消息,所有集合都获得了高达2.15的改进,对于2KB消息,所有集合都分别获得了高达2.26的改进,对于4KB消息,所有集合都获得了高达2.15的改进,对于2KB消息,所有集合都获得了高达2.26的改进。对于所有聚集,我们的SMP感知Bruck算法优于所有其他所有聚集算法,包括elan_gather(),用于512 B至8 KB消息,对于4 KB消息具有1.96的改进因子。我们的多端口Direct all-gather是16 KB到1 MB的最佳算法,对于32 KB的消息,性能优于selan_gather()1.49倍。真实的应用的实验表明,使用所提出的全聚集算法可以实现高达1.47的通信加速比。
Clusters of Symmetric Multiprocessors (SMP) are more commonplace than ever in achieving high-performance. Scientific applications running on clusters employ collective communications extensively. Shared memory communication and Remote Direct Memory Access (RDMA) over multi-rail networks are promising approaches in addressing the increasing demand on intra-node and inter-node communications, and thereby in boosting the performance of collectives in emerging multi-core SMP clusters. In this regard, this paper designs and evaluates two classes of collective communication algorithms directly at the Elan user-level over multi-rail Quadrics QsNetIIwith message striping: 1) RDMA-based traditional multi-port algorithms for gather, all-gather, and all-to-all collectives for medium to large messages, and 2) RDMA-based and SMP-aware multi-port all-gather algorithms for small to medium size messages.The multi-port RDMA-based Direct algorithm for gather and all-to-all collectives gain an improvement of up to 2.15 for 4 KB messages overelan_gather(), and up to 2.26 for 2 KB messages overelan_alltoall(), respectively. For the all-gather, our SMP-aware Bruck algorithm outperforms all other all-gather algorithms includingelan_gather()for 512 B to 8 KB messages, with a 1.96 improvement factor for 4 KB messages. Our multi-port Direct all-gather is the best algorithm for 16 KB to 1 MB, and outperformselan_gather()by a factor of 1.49 for 32 KB messages. Experimentation with real applications has shown up to 1.47 communication speedup can be achieved using the proposed all-gather algorithms.