GRAVIDY, a GPU modular, parallel direct-summation N-body integrator: Dynamics with softening

GRAVIDY, a GPU modular, parallel direct-summation N-body integrator: Dynamics with softening
复制标题

DOI:
10.1093/mnras/stx2468
复制
发表时间:
2017-02
影响因子:
4.8
通讯作者:
C. Maureira-Fredes;P. Amaro-Seoane
C. Maureira-Fredes;P. Amaro-Seoane
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
C. Maureira-Fredes;P. Amaro-Seoane

文献摘要

相似文献

天体物理学中的许多悬而未决的问题都涉及到大量粒子在重力作用下的运动。这些包括球状星团的全球演化、大质量黑洞对恒星的潮汐破坏、原行星的形成和引力辐射源的探测。引力的直接求和是一个没有解析解的复杂问题,只能用近似和数值方法来解决。为此,Hermite格式是一种广泛使用的积分方法。结合不同的数值技术和专用硬件,可以用来加快计算速度。但这些方法往往在计算上很慢,而且使用起来很麻烦。在这里,我们提出了一种新的图形处理器,直接求和$N-$体积分器从头开始,并基于该方案。该代码具有高度的模块化,允许用户容易地引入新的物理,它利用了可用的高性能计算资源,并将通过公开的、定期的更新进行维护。该代码可以在多个CPU和GPU上并行使用,具有相当大的加速优势。与单CPU版本相比,单GPU版本的运行速度大约快200倍。使用4个GPU并行运行的测试显示,与单个GPU版本相比,速度提高了约3倍。第一个版本的概念和设计针对的是能够通过一个或几个GPU卡访问传统并行CPU集群或计算节点的用户。
A wide variety of outstanding problems in astrophysics involve the motion of a large number of particles ($N\gtrsim 10^{6}$) under the force of gravity. These include the global evolution of globular clusters, tidal disruptions of stars by a massive black hole, the formation of protoplanets and the detection of sources of gravitational radiation. The direct-summation of $N$ gravitational forces is a complex problem with no analytical solution and can only be tackled with approximations and numerical methods. To this end, the Hermite scheme is a widely used integration method. With different numerical techniques and special-purpose hardware, it can be used to speed up the calculations. But these methods tend to be computationally slow and cumbersome to work with. Here we present a new GPU, direct-summation $N-$body integrator written from scratch and based on this scheme. This code has high modularity, allowing users to readily introduce new physics, it exploits available high-performance computing resources and will be maintained by public, regular updates. The code can be used in parallel on multiple CPUs and GPUs, with a considerable speed-up benefit. The single GPU version runs about 200 times faster compared to the single CPU version. A test run using 4 GPUs in parallel shows a speed up factor of about 3 as compared to the single GPU version. The conception and design of this first release is aimed at users with access to traditional parallel CPU clusters or computational nodes with one or a few GPU cards.