ConnectX-2 CORE-Direct Enabled Asynchronous Broadcast Collective Communications

ConnectX-2 CORE-Direct Enabled Asynchronous Broadcast Collective Communications
复制标题

ConnectX-2 CORE-Direct 支持异步广播集体通信

DOI:
--
复制
发表时间:
2011
期刊:
IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子:
--
通讯作者:
G. Shainer
G. Shainer
中科院分区:
--
文献类型:
--
作者:
Manjunath Gorentla Venkata;R. Graham;Joshua Ladd;Pavel Shamis;Ishai Rabinovitz;Vasily Filipov;G. Shainer

文献摘要

被引文献

相似文献

本文描述了在Cheetah集体操作框架内基于InfiniBand(IB){CORE-textit{Direct}}的阻塞和非阻塞广播操作的设计和实现。它描述了一种新的方法,完全卸载集体操作,只采用用户提供的缓冲区。对于64秩通信器,基于{CORE-textit{Direct}}的分层算法的延迟优于生产级消息传递接口(MPI)实现,对于一千字节(KB)消息,比默认的Open MPI算法好150%,比共享存储器优化的MVAPICH实现好115%,对于八兆字节(MB)消息,它好48%和64%。分别平面拓扑广播在基于轮询的通信计算测试中实现了99.9%的重叠,并且在基于等待的测试中实现了95.1%的重叠,而对于类似的基于中央处理单元(CPU)的实现,分别为92.4%和17.0%。
This paper describes the design and implementation of InfiniBand (IB) {CORE-textit{Direct}} based blocking and nonblocking broadcast operations within the Cheetah collective operation framework. It describes a novel approach that fully offloads collective operations and employs only user-supplied buffers. For a 64 rank communicator, the latency of {CORE-textit{Direct}} based hierarchical algorithm is better than production-grade Message Passing Interface (MPI) implementations, 150% better than the default Open MPI algorithm and 115% better than the shared memory optimized MVAPICH implementation for a one kilo-byte (KB) message, and for eight mega-bytes (MB) it is 48% and 64% better, respectively. Flat-topology broadcast achieves 99.9% overlap in a polling based communication-computation test, and 95.1% overlap for a wait based test, compared with 92.4% and 17.0%, respectively, for a similar Central Processing Unit (CPU) based implementation.