CU++ET: An Object Oriented Tool for Accelerating Computational Fluid Dynamics Codes using Graphical Processing Units
CU++ET: An Object Oriented Tool for Accelerating Computational Fluid Dynamics Codes using Graphical Processing Units
复制标题
CU ET:使用图形处理单元加速计算流体动力学代码的面向对象工具
DOI:
10.2514/6.2011-3222
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
D. Mavriplis
中科院分区:
文献类型:
--
作者:
J. Sitaraman;D. Mavriplis
The application of graphical processing units (GPU) to solve partial dierential equations is gaining popularity with the advent of improved computer hardware. Various lower level interfaces exist that allow the user to access GPU specic functions. One such interface is NVIDIA’s Compute Unied Device Architecture (CUDA) library. CUDA has been applied previously to solve the Three-Dimensional Euler equations, and a speed-up of the order of 500 has been reported in literature using multiple GPU units. However, porting existing codes to run on the GPU requires the user to write kernels that execute on multiple cores, in the form of Single Instruction Multiple Data (SIMD). In the present work, a higher level framework has been developed that uses object oriented programming techniques available in C++ such as polymorphism, operator overloading, and template meta programming. Using this approach, CUDA kernels can be generated automatically during compile time. Briey, CU++ET allows a code developer with just C/C++ knowledge to write computer programs that will execute on the GPU without any knowledge of specic programming techniques in CUDA. It allows the user to reuse existing C/C++ CFD codes with minimal changes. This approach is tremendously benecial for CFD code development because it mitigates the necessity of creating hundreds of GPU kernels for various purposes. In its current form, CU++ET provides a framework for parallel array arithmetic, simplied data structures to interface with the GPU, and smart array indexing. Using this framework, a higher-order 3D Euler solver (ARC3D-GPU) has been developed with a performance improvement of about 70x on a single GPU compared to traditional FORTRAN/CPU execution. An implementation of heterogeneous parallelism, i.e., utilizing multiple GPUs to simultaneously process a partitioned grid system with communication at the interfaces using MPI has been developed and tested. An unstructured version of CU++ET is also demonstrated with its application towards solving the incompressible Navier-Stokes equations.