Exploiting temporal parallelism in particle-based incompressive fluid simulation on FPGA
Exploiting temporal parallelism in particle-based incompressive fluid simulation on FPGA
复制标题
在 FPGA 上基于粒子的非压缩流体模拟中利用时间并行性
DOI:
10.1109/candar51075.2020.00034
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Miyajima Takaaki
中科院分区:
文献类型:
--
作者:
Orsztynowicz Manfred;Amano Hideharu;Kubota Kenichi;Miyajima Takaaki
While the semiconductor manufacturing process is shrinking, the Bytes per Flop (B/F) ratio on recent machines is becoming lower. Particle-based computational fluid dynamics (CFD) methods such as Moving Particle Simulation (MPS) require a higher B/F ratio than that of stencil-based CFD methods. Techniques to reduce the B/F ratio by exploiting temporal parallelism is becoming popular in stencil-based CFD methods on CPU and GPU. It is also reported that a technique to combine temporal blocking with stencil buffer is suitable for FPGA and can outperform CPU and GPU. On the other hand, it has been considered that temporal parallelism cannot be exploited in the particle-based CFD methods. This is because the number of particles in each bucket, a three-dimensional grid covering a computational domain, changes every time-step. In this paper, we propose a technique to exploit temporal parallelism in MPS method, a particle-based CFD method for incompressive fluid. The key idea is that the buckets in MPS can be considered as stencils in stencil-based CFD. This is because the maximum number of particles in a bucket can be assumed empirically in the case of an incompressible fluid. To the best of our knowledge, this is the first research which exploits temporal parallelism in the particle-based incompressible fluid method. We implemented the proposed technique with a degree of temporal parallelism of three. We also optimized it on Intel Arria10 FPGA in Intel HLS, and measured the performance and resource consumption. The result shows that the optimized implementation with a degree of temporal parallelism of three achieved 2.1 times speedup compared with implementation without exploiting temporal parallelism on CPU.
登录
查看更多内容
DOI:
10.1109/icis.2016.7550742
发表时间:
2016-08
期刊:
2016 IEEE/ACIS 15th International Conference on Computer and Information Science (ICIS)
影响因子:
--
作者:
Hasitha Muthumala Waidyasooriya;M. Hariyama
通讯作者:
Hasitha Muthumala Waidyasooriya;M. Hariyama
DOI:
10.1016/j.procs.2015.05.315
发表时间:
2015
期刊:
2013 IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子:
--
作者:
T. Muranushi;J. Makino
通讯作者:
J. Makino
DOI:
--
发表时间:
2012
期刊:
影响因子:
--
作者:
S. Koshizuka;M. Oochi; K. Shibata
通讯作者:
K. Shibata
影响因子:
1.2
作者:
Koshizuka, S;Oka, Y
通讯作者:
Oka, Y
DOI:
--
发表时间:
2011
期刊:
Proc.19th Annual IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM2011)
影响因子:
--
作者:
Kentaro Sano;Yoshiaki Hatsuda;Satoru Yamamoto
通讯作者:
Satoru Yamamoto