MPI vs Fortran coarrays beyond 100k cores: 3D cellular automata

MPI vs Fortran coarrays beyond 100k cores: 3D cellular automata
复制标题

DOI:
10.1016/j.parco.2019.03.002
复制
发表时间:
2019-05-01
期刊:
影响因子:
1.4
通讯作者:
Cebamanos, Luis
Cebamanos, Luis
中科院分区:
计算机科学4区
文献类型:
--
作者:
Shterenlikht, Anton;Cebamanos, Luis

文献摘要

被引文献

相似文献

由于熟悉的Fortran语法、单边通信和在编译器中的实现,Fortran共数组是MPI的一个有吸引力的替代方案。在这项工作中,使用由作者开发的CASUP CA库https://cgpack.sourceforge.io,构建的元胞自动机(CA)3D伊辛磁化微型应用程序,将共阵列的缩放与MPI进行了比较。使用MPI_ALLREDUCE和Fortran 2018 co_sum集合体计算了Ising能量和磁化强度。这项工作是在Archer(Cray XC30)上完成的,直到机器的全部容量:109,056个核心。乒乓延迟和带宽结果与MPI和消息大小从1B到几MB的共阵列非常相似。MPI Halo Exchange(HX)比协阵列HX的可扩展性更好,这令人惊讶,因为这两种算法都使用成对通信:MPI IRECV/ISEND/WAITALL与Fortran同步图像。将OpenMP添加到MPI或共阵列会导致更差的二级缓存命中率,并且在所有情况下都会降低性能,即使排除了NUMA影响。这很可能是因为CA算法在规模上受网络限制。非常激进的缓存和过程间优化不会带来任何性能提升,这进一步证明了这一点。采样和跟踪分析表明,在所有小应用程序中,计算负载平衡良好,但通信不平衡,这表明MPI和协数组之间的性能差异可能是并行库(MPICH2与libpgas)和Cray硬件特定库(ugni与DMAPP)造成的。总体而言,结果看起来很有希望用于10万核以上的共阵列。然而,需要进一步的共阵列优化以缩小共阵列与MPI之间的性能差距。(C)2019爱思唯尔B.V.保留所有权利。
Fortran coarrays are an attractive alternative to MPI due to a familiar Fortran syntax, single sided communications and implementation in the compiler. Scaling of coarrays is compared in this work to MPI, using cellular automata (CA) 3D Ising magnetisation miniapps, built with the CASUP CA library, https://cgpack.sourceforge.io, developed by the authors. Ising energy and magnetisation were calculated with MPI_ALLREDUCE and Fortran 2018 co_sum collectives. The work was done on ARCHER (Cray XC30) up to the full machine capacity: 109,056 cores. Ping-pong latency and bandwidth results are very similar with MPI and with coarrays for message sizes from 1B to several MB. MPI halo exchange (HX) scaled better than coarray HX, which is surprising because both algorithms use pair-wise communications: MPI IRECV/ISEND/WAITALL vs Fortran sync images. Adding OpenMP to MPI or to coarrays resulted in worse L2 cache hit ratio, and lower performance in all cases, even though the NUMA effects were ruled out. This is likely because the CA algorithm is network bound at scale. This is further evidenced by the fact that very aggressive cache and inter-procedural optimisations lead to no performance gain. The sampling and tracing analysis shows good load balancing in compute in all miniapps, but imbalance in communication, indicating that the difference in performance between MPI and coarrays is likely due to parallel libraries (MPICH2 vs libpgas) and the Cray hardware specific libraries (uGNI vs DMAPP). Overall, the results look promising for coarray use beyond 100k cores. However, further coarray optimisation is needed to narrow the performance gap between coarrays and MPI. (C) 2019 Elsevier B.V. All rights reserved.