Large Stencil Operations for GPU-based 3-D Acoustics Simulations

Large Stencil Operations for GPU-based 3-D Acoustics Simulations
复制标题

基于 GPU 的 3D 声学仿真的大型模板操作

DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
S. Bilbao
S. Bilbao
中科院分区:
--
文献类型:
--
作者:
B. Hamilton;Craig J. Webb;A. Gray;S. Bilbao

文献摘要

被引文献

相似文献

在执行声学模拟时,模板操作通常是一个关键组成部分,对于该模拟,实施的特定选择可以对准确性和计算性能产生重大影响。有限的差异方案,使用不同形状和尺寸的模具(使用NVIDIA K20 GPU的尺寸为7到450点)。随着模板尺寸的增加,通过仅考虑涉及的计算操作的数量,计算时间的增加小于自然预期的时间,因为在整个GPU内存体系结构中,性能是由数据传输时间确定的。在空间中紧凑的模具的主要是由于在K20上有效地使用了仅读取数据(纹理)缓存,并且具有标准高阶的性能模具是由于记忆带宽的使用增加,在本研究中还要弥补较低的高速播放率发现,通过有效利用GTX 660TI GPU(其计算性能通常低于GTX 670),类似或更好的性能可以是在不使用共享内存的情况下实现。
Stencil operations are often a key component when performing acoustics simulations, for which the specific choice of implementation can have a significant effect on both accuracy and computational performance. This paper presents a detailed investigation of computational performance for GPU-based stencil operations in two-step finite difference schemes, using stencils of varying shape and size (ranging from seven to more than 450 points in size). Using an Nvidia K20 GPU, it is found that as the stencil size increases, compute times increase less than that naively expected by considering only the number of computational operations involved, because performance is instead determined by data transfer times throughout the GPU memory architecture. With regards to the effects of stencil shape, performance obtained with stencils that are compact in space is mainly due to efficient use of the read-only data (texture) cache on the K20, and performance obtained with standard high-order stencils is due to increased memory bandwidth usage, compensating for lower cache hit rates. Also in this study, a brief comparison is made with performance results from a related, recent study that used a shared memory approach on a GTX 670 GPU device. It is found that by making efficient use of a GTX 660Ti GPU—whose computational performance is generally lower than that of a GTX 670—similar or better performance to those results can be achieved without the use of shared memory.