Manycore challenge in particle-in-cell simulation: How to exploit 1 TFlops peak performance for simulation codes with irregular computation

Manycore challenge in particle-in-cell simulation: How to exploit 1 TFlops peak performance for simulation codes with irregular computation
复制标题

DOI:
10.1016/j.compeleceng.2015.03.010
复制
发表时间:
2015-08
期刊:
Comput. Electr. Eng.
影响因子:
--
通讯作者:
H. Nakashima
H. Nakashima
中科院分区:
其他
文献类型:
--
作者:
H. Nakashima

文献摘要

相似文献

本文讨论了后Peta和Exascale时代的挑战,特别是普通(即,非GPU类型)CPU核心。虽然像英特尔至强融核这样的处理器为我们提供了TFlops级的计算能力,并可能引导我们进行Exascale计算,但由于其高性能的来源,即大规模多线程和广泛的SIMD机制,充分利用其潜力远非易事。事实上,在三层并行中,即节点间,节点内和内核内,我们发现它们的顺序并不代表HPC编程的韧性,但顺序应该颠倒。我们的粒子在细胞等离子体模拟代码的案例研究支持我们的观察,揭示了一个简单的移植现有的代码到至强融核是不可行的,从性能的角度来看,我们必须作出重大改变的代码结构,使其符合处理器的功能。然而,该研究还证实,重新编码的努力是很好的回报,实现了良好的单节点性能高于从Cray XE6的四个双插槽节点上执行获得。
This paper discusses the challenge in post-Peta and Exascale era especially that brought by manycore processors of ordinary (i.e., non-GPU type) CPU cores. Though such a processor like Intel Xeon Phi gives us TFlops-class computational power and may lead us to Exascale computing, full exploitation of its potential is far from an easy job due to its source of high performance, namely a large scale multithreading and a wide SIMD mechanism. In fact, in the three-tier parallelism namely inter-node, intra-node and intra-core ones, we found their order does not represent the toughness in HPC programming but the order should be reversed to do that. Our case study with a particle-in-cell plasma simulation code supports our observation revealing that a simple porting of an existing code to Xeon Phi is infeasible from the viewpoint of performance and we have to make a significant change of the code structure so that it conforms with the features of the processor. However the study also confirms that the recoding effort is well rewarded achieving a good single-node performance higher than that obtained from an execution on four dual-socket nodes of Cray XE6.