Many-Core Acceleration of a Discrete Ordinates Transport Mini-App at Extreme Scale

Many-Core Acceleration of a Discrete Ordinates Transport Mini-App at Extreme Scale
复制标题

超大规模离散坐标传输小应用程序的多核加速

DOI:
10.1007/978-3-319-41321-1_22
复制
发表时间:
2016
期刊:
The New England journal of medicine
影响因子:
--
通讯作者:
W. Gaudin
W. Gaudin
中科院分区:
--
文献类型:
--
作者:
Tom Deakin;Simon McIntosh;W. Gaudin

文献摘要

被引文献

相似文献

时间相关的确定性离散坐标传输码是一类重要的应用,它为大型众核系统提供了重大挑战。其中一个挑战是求解步骤所需的大内存容量,这要求我们有一个可扩展的解决方案,以便有足够的节点级内存来存储所有数据。在我们之前的工作中,我们演示了第一个实现,该实现显示出使用GPU的单节点求解的显著性能优势。在本文中,我们将我们的工作扩展到大的问题,并证明了我们的解决方案的可扩展性上的两个Petascale基于GPU的超级计算机:泰坦在橡树岭和Piz Daint在CSCS。我们的研究结果表明,我们改进的节点级并行计划的规模以及在大型系统中使用久经考验的KBA域分解技术时,以前的方法。我们验证我们的结果对改进的性能模型,预测运行时的主要'扫描'例程在不同的硬件,包括CPU或GPU上运行。
Time-dependent deterministic discrete ordinates transport codes are an important class of application which provide significant challenges for large, many-core systems. One such challenge is the large memory capacity needed by the solve step, which requires us to have a scalable solution in order to have enough node-level memory to store all the data. In our previous work, we demonstrated the first implementation which showed a significant performance benefit for single node solves using GPUs. In this paper we extend our work to large problems and demonstrate the scalability of our solution on two Petascale GPU-based supercomputers: Titan at Oak Ridge and Piz Daint at CSCS. Our results show that our improved node-level parallelism scheme scales just as well across large systems as previous approaches when using the tried and tested KBA domain decomposition technique. We validate our results against an improved performance model which predicts the runtime of the main ‘sweep’ routine when running on different hardware, including CPUs or GPUs.