An Optimized Multicolor Point-Implicit Solver for Unstructured Grid Applications on Graphics Processing Units

An Optimized Multicolor Point-Implicit Solver for Unstructured Grid Applications on Graphics Processing Units
复制标题

用于图形处理单元上非结构化网格应用的优化多色点隐式求解器

DOI:
--
复制
发表时间:
2016
期刊:
Workshop on Irregular Applications: Architectures and Algorithms
影响因子:
--
通讯作者:
D. Hammond
D. Hammond
中科院分区:
--
文献类型:
--
作者:
M. Zubair;E. Nielsen;Justin Luitjens;D. Hammond

文献摘要

被引文献

相似文献

在计算流体力学领域中,为了适应几何复杂性,通常使用非结构网格方法来求解Navier-Stokes方程。这种空间离散化的隐式求解方法通常需要频繁地求解大型紧耦合块稀疏线性方程组。当前工作中使用的多色点隐式求解器通常需要整个应用程序运行时间的很大一部分。在这项工作中,提出了一种高效的图形处理单元求解器的实现。有几个因素对在这种环境下实现有效实施提出了独特的挑战。这些问题包括不同内核调用中可用的可变并行度、间接内存访问模式、低算术强度以及支持可变块大小的要求。在这项工作中,求解器被重新定义为使用标准的稀疏和稠密的基本线性代数子程序(BLAS)函数。然而,实验表明,对于实际模拟中遇到的矩阵,现有CUDA库中可用的BLAS函数的性能不是最优的。相反,开发了这些函数的优化版本。根据块大小的不同,新实施的性能比现有CUDA库函数高出7倍。
In the field of computational fluid dynamics, the Navier-Stokes equations are often solved using an unstructured-grid approach to accommodate geometric complexity. Implicit solution methodologies for such spatial discretizations generally require frequent solution of large tightly-coupled systems of block-sparse linear equations. The multicolor point-implicit solver used in the current work typically requires a significant fraction of the overall application run time. In this work, an efficient implementation of the solver for graphics processing units is proposed. Several factors present unique challenges to achieving an efficient implementation in this environment. These include the variable amount of parallelism available in different kernel calls, indirect memory access patterns, low arithmetic intensity, and the requirement to support variable block sizes. In this work, the solver is reformulated to use standard sparse and dense Basic Linear Algebra Subprograms (BLAS) functions. However, experiments show that the performance of the BLAS functions available in existing CUDA libraries is suboptimal for matrices representative of those encountered in actual simulations. Instead, optimized versions of these functions are developed. Depending on block size, the new implementations show performance gains of up to 7× over the existing CUDA library functions.