Multi-core acceleration of chemical kinetics for simulation and prediction

Multi-core acceleration of chemical kinetics for simulation and prediction
复制标题

DOI:
10.1145/1654059.1654067
复制
发表时间:
2009-11
期刊:
Proceedings of the Conference on High Performance Computing Networking, Storage and Analysis
影响因子:
--
通讯作者:
J. C. Linford;J. Michalakes;Manish Vachharajani;Adrian Sandu
J. C. Linford;J. Michalakes;Manish Vachharajani;Adrian Sandu
中科院分区:
其他
文献类型:
--
作者:
J. C. Linford;J. Michalakes;Manish Vachharajani;Adrian Sandu

文献摘要

被引文献

相似文献

这项工作实现了三个多核平台上大型社区大气模型的计算昂贵的化学动力学内核:使用CUDA,细胞宽带发动机和英特尔四核Xeon CPU的NVIDIA GPU。提出了对每个平台的比较性能分析,以粗糙和细网格的双重精度和单个精度进行。平台特异性的设计和优化以机制敏捷的方式讨论,允许优化许多化学机制。讨论了用于SIMD架构的三阶段Rosenbrock求解器的实现。当用作动力学预处理器中的模板机制时,多核实现可以在各种多核平台上自动优化和移植许多化学机制。与八个Xeon芯相比,单个精度的加速度为5.5倍,双精度为2.7倍。与串行实现相比,最大观察到的速度在单个精度中为41.1倍。
This work implements a computationally expensive chemical kinetics kernel from a large-scale community atmospheric model on three multi-core platforms: NVIDIA GPUs using CUDA, the Cell Broadband Engine, and Intel Quad-Core Xeon CPUs. A comparative performance analysis for each platform in double and single precision on coarse and fine grids is presented. Platform-specific design and optimization is discussed in a mechanism-agnostic way, permitting the optimization of many chemical mechanisms. The implementation of a three-stage Rosenbrock solver for SIMD architectures is discussed. When used as a template mechanism in the the Kinetic PreProcessor, the multi-core implementation enables the automatic optimization and porting of many chemical mechanisms on a variety of multi-core platforms. Speedups of 5.5x in single precision and 2.7x in double precision are observed when compared to eight Xeon cores. Compared to the serial implementation, the maximum observed speedup is 41.1x in single precision.