42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence

42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence
复制标题

DOI:
10.1145/1654059.1654123
复制
发表时间:
2009-11
期刊:
Proceedings of the Conference on High Performance Computing Networking, Storage and Analysis
影响因子:
--
通讯作者:
T. Hamada;T. Narumi;Rio Yokota;K. Yasuoka;Keigo Nitadori;M. Taiji
T. Hamada;T. Narumi;Rio Yokota;K. Yasuoka;Keigo Nitadori;M. Taiji
中科院分区:
其他
文献类型:
--
作者:
T. Hamada;T. Narumi;Rio Yokota;K. Yasuoka;Keigo Nitadori;M. Taiji

文献摘要

被引文献

相似文献

作为2009年戈登贝尔价格/性能奖的参赛作品,我们提出了两种不同的层次N体模拟的结果,集群的256个图形处理单元(GPU)。与许多以前的N体模拟的GPU上的规模为O(N2),本方法计算O(N log N)的树码和O(N)的快速多极子方法(FMM)的GPU上前所未有的效率。我们证明了我们的方法的性能,选择一个标准的应用程序-重力N体模拟-和一个非标准的应用程序-模拟湍流使用涡粒子。使用具有1,608,044,129个粒子的树码进行的重力模拟显示出42.15 TFlops的持续性能。使用具有16,777,216个粒子的周期性FMM对均匀各向同性湍流的涡粒子模拟显示出20.2 TFlops的持续性能。硬件的总成本为228,912美元。重力模拟的最大校正性能为28.1TFlops,这导致124 MFlops/$的性价比。这种校正是通过基于最有效的CPU算法对触发器进行计数来执行的。GPU实现和参数差异产生的任何额外的Flop都不包括在124 MFlops/$中。
As an entry for the 2009 Gordon Bell price/performance prize, we present the results of two different hierarchical N-body simulations on a cluster of 256 graphics processing units (GPUs). Unlike many previous N-body simulations on GPUs that scale as O(N2), the present method calculates the O(N log N) treecode and O(N) fast multipole method (FMM) on the GPUs with unprecedented efficiency. We demonstrate the performance of our method by choosing one standard application -a gravitational N-body simulation- and one non-standard application -simulation of turbulence using vortex particles. The gravitational simulation using the treecode with 1,608,044,129 particles showed a sustained performance of 42.15 TFlops. The vortex particle simulation of homogeneous isotropic turbulence using the periodic FMM with 16,777,216 particles showed a sustained performance of 20.2 TFlops. The overall cost of the hardware was 228,912 dollars. The maximum corrected performance is 28.1TFlops for the gravitational simulation, which results in a cost performance of 124 MFlops/$. This correction is performed by counting the Flops based on the most efficient CPU algorithm. Any extra Flops that arise from the GPU implementation and parameter differences are not included in the 124 MFlops/$.