Numerical ocean modeling and simulation with CUDA

Numerical ocean modeling and simulation with CUDA
复制标题

DOI:
10.23919/oceans.2011.6107199
复制
发表时间:
2011-12
期刊:
OCEANS'11 MTS/IEEE KONA
影响因子:
--
通讯作者:
J. Mak;P. Choboter;C. Lupo
J. Mak;P. Choboter;C. Lupo
中科院分区:
其他
文献类型:
--
作者:
J. Mak;P. Choboter;C. Lupo

文献摘要

被引文献

相似文献

ROMS是使用有限差分网格和时间步进对海洋区域进行建模和模拟的软件。由于软件的计算密集型性质,ROMS模拟可能需要数小时至数天才能完成。因此,模拟的大小和分辨率受到现代计算硬件的性能限制的约束。为了解决这些问题,现有的ROMS代码可以与OpenMP或MPI并行运行。在这项工作中,我们实现了一个新的并行化的ROM上的图形处理单元(GPU)使用CUDA Fortran。我们利用现代GPU提供的大规模并行性,以更低的成本和更少的功耗获得性能优势。为了测试我们的实现,我们基准与理想的海洋条件,以及真实的数据收集从沿海沃茨附近的中部加州。我们的实现产生了高达8倍的串行实现和2.5倍以上的OpenMP实现的加速比,同时表现出可比的性能MPI实现。
ROMS is software that models and simulates an ocean region using a finite difference grid and time stepping. ROMS simulations can take from hours to days to complete due to the compute-intensive nature of the software. As a result, the size and resolution of simulations are constrained by the performance limitations of modern computing hardware. To address these issues, the existing ROMS code can be run in parallel with either OpenMP or MPI. In this work, we implement a new parallelization of ROMS on a graphics processing unit (GPU) using CUDA Fortran. We exploit the massive parallelism offered by modern GPUs to gain a performance benefit at a lower cost and with less power. To test our implementation, we benchmark with idealistic marine conditions as well as real data collected from coastal waters near central California. Our implementation yields a speedup of up to 8x over a serial implementation and 2.5x over an OpenMP implementation, while demonstrating comparable performance to a MPI implementation.