Performance Measurements of the 3D FFT on the Blue Gene/L Supercomputer

Performance Measurements of the 3D FFT on the Blue Gene/L Supercomputer
复制标题

Blue Gene/L 超级计算机上 3D FFT 的性能测量

DOI:
10.1007/11549468_87
复制
发表时间:
2005
期刊:
--
影响因子:
--
通讯作者:
R. Germain
R. Germain
中科院分区:
--
文献类型:
--
作者:
M. Eleftheriou;B. Fitch;A. Rayshubskiy;T. Ward;R. Germain

文献摘要

被引文献

相似文献

本文介绍了一个通信密集型的内核,复杂的数据三维FFT,运行在蓝色基因/L架构的性能特点。对体积FFT算法的两种实现进行了表征,一种是使用优化的集体全对全操作建立在MPI库上[2],另一种是建立在Blue Gene/L高级诊断环境(BG/L ADE)的低级系统编程接口(SPI)上[17]。我们将当前结果与使用参考MPI实现(MPICH 2移植到BG/L,未优化集合)和FFTW库2.1.5版本的端口[14]获得的结果进行比较。在Blue Gene/L原型上的性能实验表明,我们的两种实现都具有很好的扩展性,并且当前基于MPI的实现在2048个节点上对大小为128 × 128 × 128的3D FFT的加速比为730。此外,对于2048个节点上的128× 128× 128复FFT,体积FFT的性能优于FFTW端口8倍。
This paper presents performance characteristics of a communications-intensive kernel, the complex data 3D FFT, running on the Blue Gene/L architecture. Two implementations of the volumetric FFT algorithm were characterized, one built on the MPI library using an optimized collective all-to-all operation [2] and another built on a low-level System Programming Interface (SPI) of the Blue Gene/L Advanced Diagnostics Environment (BG/L ADE) [17]. We compare the current results to those obtained using a reference MPI implementation (MPICH2 ported to BG/L with unoptimized collectives) and to a port of version 2.1.5 the FFTW library [14]. Performance experiments on the Blue Gene/L prototype indicate that both of our implementations scale well and the current MPI-based implementation shows a speedup of 730 on 2048 nodes for 3D FFTs of size 128 × 128 × 128. Moreover, the volumetric FFT outperforms FFTW port by a factor 8 for a 128× 128× 128 complex FFT on 2048 nodes.