Accelerating multi-dimensional population balance model simulations via a highly scalable framework using GPUs

Accelerating multi-dimensional population balance model simulations via a highly scalable framework using GPUs
复制标题

DOI:
10.1016/j.compchemeng.2020.106935
复制
发表时间:
2020-09
期刊:
Comput. Chem. Eng.
影响因子:
--
通讯作者:
Chaitanya Sampat;Y. Baranwal;R. Ramachandran
Chaitanya Sampat;Y. Baranwal;R. Ramachandran
中科院分区:
其他
文献类型:
--
作者:
Chaitanya Sampat;Y. Baranwal;R. Ramachandran

文献摘要

被引文献

相似文献

使用CPU的高维PBM的解决方案通常在计算上是困难的。这项研究致力于开发一种可扩展的算法,通过一个GPU框架来并行化PBM中的嵌套循环。开发的PBM是独一无二的,因为它适应了问题的大小,并相应地使用了GPU内核。该算法在使用CUDA®和C/C++编写时针对NVIDIA®GPU进行了并行化。这类算法的主要瓶颈是CPU和GPU之间的通信时间。在我们的研究中,通信时间只占总运行时间的不到1%,与串行CPU代码相比,最大加速比达到了约12%。与台式计算机上的多核配置相比,GPU PBM实现了大约两倍的加速。据报道,各种CPU和GPU架构和配置的速度也有所提高。
The solution of high-dimensional PBMs using CPUs are often computationally intractable. This study focuses on the development of a scalable algorithm to parallelize the nested loops inside the PBM via a GPU framework. The developed PBM is unique since it adapts to the size of the problem and uses the GPU cores accordingly. This algorithm was parallelized for NVIDIA® GPUs as it was written in CUDA® and C/C++. The major bottleneck of such algorithms is the communication time between the CPU and the GPU. In our studies, communication time contributed to less than 1% of the total run time and a maximum speedup of about 12 over the serial CPU code was achieved. The GPU PBM achieved a speedup of about two times compared to the PBM’s multi-core configuration on a desktop computer. The speed improvements are also reported for various CPU and GPU architectures and configurations.