Implementation and performance of FDPS: a framework for developing parallel particle simulation codes

Implementation and performance of FDPS: a framework for developing parallel particle simulation codes
复制标题

DOI:
10.1093/pasj/psw053
复制
发表时间:
2016-08-01
影响因子:
2.3
通讯作者:
Makino, Junichiro
Makino, Junichiro
中科院分区:
物理与天体物理4区
文献类型:
--
作者:
Iwasawa, Masaki;Tanikawa, Ataru;Makino, Junichiro

文献摘要

被引文献

相似文献

介绍了粒子模拟器开发框架FDPS的基本思想、实现方法、实测性能和性能模型。FDPS是一个应用程序开发框架,它帮助研究人员使用粒子方法为大规模分布式内存并行超级计算机开发模拟程序。基于粒子的分布存储并行计算机仿真程序需要进行区域分解、交换不在每个计算节点域内的粒子以及收集其他节点上的粒子信息,这些都是进行交互计算所必需的。此外,即使不使用分布式存储器并行计算机,为了减少计算量,在长距离相互作用的情况下,也应该使用诸如Barnes-Hut树算法或快速多极方法之类的算法。对于短程相互作用,需要一些方法将计算限制在相邻粒子。FDPS提供了所有这些功能,这些功能是高效并行执行基于粒子的模拟所必需的“模板”,这些功能独立于粒子的实际数据结构和粒子-粒子相互作用的功能形式。通过使用FDPS,研究人员可以编写他们的程序所需的工作量来编写一个简单的,顺序的和未优化的程序的O(N-2)计算成本,但程序,一旦编译与FDPS,将有效地运行在大规模并行超级计算机。一个简单的引力N体程序可以写在大约120行。我们报告这些程序的实际性能和性能模型。弱标度性能是非常好的,几乎线性的速度提高,直到K计算机的整个系统。每个时间步的最小计算时间在30 ms(N = 10(7))至300 ms(N = 10(9))的范围内。目前,这些方法受到相互作用计算所需的区域分解计算和通信时间的限制。我们讨论如何克服这些瓶颈。
We present the basic idea, implementation, measured performance, and performance model of FDPS (Framework for Developing Particle Simulators). FDPS is an application-development framework which helps researchers to develop simulation programs using particle methods for large-scale distributed-memory parallel supercomputers. A particle-based simulation program for distributed-memory parallel computers needs to perform domain decomposition, exchange of particles which are not in the domain of each computing node, and gathering of the particle information in other nodes which are necessary for interaction calculation. Also, even if distributed-memory parallel computers are not used, in order to reduce the amount of computation, algorithms such as the Barnes-Hut tree algorithm or the Fast Multipole Method should be used in the case of long-range interactions. For short-range interactions, some methods to limit the calculation to neighbor particles are required. FDPS provides all of these functions which are necessary for efficient parallel execution of particle-based simulations as "templates," which are independent of the actual data structure of particles and the functional form of the particle-particle interaction. By using FDPS, researchers can write their programs with the amount of work necessary to write a simple, sequential and unoptimized program of O(N-2) calculation cost, and yet the program, once compiled with FDPS, will run efficiently on large-scale parallel supercomputers. A simple gravitational N-body program can be written in around 120 lines. We report the actual performance of these programs and the performance model. The weak scaling performance is very good, and almost linear speed-up was obtained for up to the full system of the K computer. The minimum calculation time per timestep is in the range of 30 ms (N = 10(7)) to 300 ms (N = 10(9)). These are currently limited by the time for the calculation of the domain decomposition and communication necessary for the interaction calculation. We discuss how we can overcome these bottlenecks.