CU++ET: An Object Oriented Tool for Accelerating Computational Fluid Dynamics Codes using Graphical Processing Units

CU++ET: An Object Oriented Tool for Accelerating Computational Fluid Dynamics Codes using Graphical Processing Units
复制标题

CU ET:使用图形处理单元加速计算流体动力学代码的面向对象工具

DOI:
10.2514/6.2011-3222
复制
发表时间:
2011
期刊:
Comput. Phys. Commun.
影响因子:
--
通讯作者:
D. Mavriplis
D. Mavriplis
中科院分区:
--
文献类型:
--
作者:
J. Sitaraman;D. Mavriplis

文献摘要

被引文献

相似文献

随着计算机硬件技术的发展,图形处理器(GPU)在求解偏微分方程中的应用越来越广泛。存在各种较低级别的接口,允许用户访问GPU专用功能。其中一个接口是NVIDIA的计算统一设备架构(CUDA)库。CUDA以前已经被应用于求解三维欧拉方程,并且在文献中已经报道了使用多个GPU单元的500级的加速。然而,将现有代码移植到GPU上运行需要用户以单指令多数据(SIMD)的形式编写在多个内核上执行的内核。在目前的工作中,已经开发出一个更高层次的框架,使用面向对象的编程技术,在C++,如多态性,运算符重载,和模板Meta编程。使用这种方法,CUDA内核可以在编译时自动生成。布里埃,CU++ET允许代码开发人员只有C/C++知识编写将在GPU上执行的计算机程序,而不需要任何CUDA中特定编程技术的知识。它允许用户以最小的更改重用现有的C/C++ CFD代码。这种方法对于CFD代码开发非常有益,因为它减轻了为各种目的创建数百个GPU内核的必要性。CU++ET提供了一个并行数组运算的框架,简化了与GPU接口的数据结构,以及智能数组索引。使用这个框架,高阶3D欧拉求解器(ARC 3D-GPU)已开发出一个单一的GPU上的性能提高了约70倍,相比传统的FORTRAN/CPU执行。异构并行的实现,即,利用多个GPU同时处理分区网格系统,并在接口处使用MPI进行通信。一个非结构化版本的CU++ET也展示了其对解决不可压Navier-Stokes方程的应用。
The application of graphical processing units (GPU) to solve partial dierential equations is gaining popularity with the advent of improved computer hardware. Various lower level interfaces exist that allow the user to access GPU specic functions. One such interface is NVIDIA’s Compute Unied Device Architecture (CUDA) library. CUDA has been applied previously to solve the Three-Dimensional Euler equations, and a speed-up of the order of 500 has been reported in literature using multiple GPU units. However, porting existing codes to run on the GPU requires the user to write kernels that execute on multiple cores, in the form of Single Instruction Multiple Data (SIMD). In the present work, a higher level framework has been developed that uses object oriented programming techniques available in C++ such as polymorphism, operator overloading, and template meta programming. Using this approach, CUDA kernels can be generated automatically during compile time. Briey, CU++ET allows a code developer with just C/C++ knowledge to write computer programs that will execute on the GPU without any knowledge of specic programming techniques in CUDA. It allows the user to reuse existing C/C++ CFD codes with minimal changes. This approach is tremendously benecial for CFD code development because it mitigates the necessity of creating hundreds of GPU kernels for various purposes. In its current form, CU++ET provides a framework for parallel array arithmetic, simplied data structures to interface with the GPU, and smart array indexing. Using this framework, a higher-order 3D Euler solver (ARC3D-GPU) has been developed with a performance improvement of about 70x on a single GPU compared to traditional FORTRAN/CPU execution. An implementation of heterogeneous parallelism, i.e., utilizing multiple GPUs to simultaneously process a partitioned grid system with communication at the interfaces using MPI has been developed and tested. An unstructured version of CU++ET is also demonstrated with its application towards solving the incompressible Navier-Stokes equations.