Benefits of Cross Memory Attach for MPI libraries on HPC Clusters

Benefits of Cross Memory Attach for MPI libraries on HPC Clusters
复制标题

HPC 集群上 MPI 库的跨内存附加的优势

DOI:
--
复制
发表时间:
2014
期刊:
Extreme Science and Engineering Discovery Environment
影响因子:
--
通讯作者:
Jérôme Vienne
Jérôme Vienne
中科院分区:
--
文献类型:
--
作者:
Jérôme Vienne

文献摘要

被引文献

相似文献

随着现代集群中每个节点的核心数量不断增加,节点内通信的有效实施对于应用程序性能至关重要。 MPI 库通常使用共享内存机制在节点内部进行通信,不幸的是这种方法对于大消息有一些限制。 Linux 内核 3.2 版本引入了跨内存附加 (CMA),这是一种改进同一节点内 MPI 进程之间通信的机制。但是,由于默认情况下在支持该功能的 MPI 库中并未启用此功能,因此 HPC 管理员可能会将其禁用,从而导致用户丧失性能优势。在本文中,我们解释了如何使用 CMA,并使用微基准测试和 NAS 并行基准测试 (NPB) 来评估 CMA,这是一组常用于评估并行系统的应用程序。 我们的性能评估表明,对于大型消息,CMA 的性能优于共享内存性能。微基准级别评估表明,CMA 可以将性能提高四倍。借助 NPB,我们发现 FT 的总执行时间提高了 24.75%,IS 的总执行时间提高了 24.08%。
With the number of cores per node increasing in modern clusters, an efficient implementation of intra-node communications is critical for application performance. MPI libraries generally use shared memory mechanisms for communication inside the node, unfortunately this approach has some limitations for large messages. The release of Linux kernel 3.2 introduced Cross Memory Attach (CMA) which is a mechanism to improve the communication between MPI processes inside the same node. But, as this feature is not enabled by default inside MPI libraries supporting it, it could be left disabled by HPC administrators which leads to a loss of performance benefits to users. In this paper, we explain how to use CMA and present an evaluation of CMA using micro-benchmarks and NAS parallel benchmarks (NPB) which are a set of applications commonly used to evaluate parallel systems. Our performance evaluation reveals that CMA outperforms shared memory performance for large messages. Micro-benchmark level evaluations show that CMA can enhance the performance by as much as a factor of four. With NPB, we see up to 24.75% improvement in total execution time for FT and up to 24.08% for IS.