Exploiting temporal parallelism in particle-based incompressive fluid simulation on FPGA

Exploiting temporal parallelism in particle-based incompressive fluid simulation on FPGA
复制标题

在 FPGA 上基于粒子的非压缩流体模拟中利用时间并行性

DOI:
10.1109/candar51075.2020.00034
复制
发表时间:
2020
期刊:
2020 Eighth International Symposium on Computing and Networking (CANDAR)
影响因子:
--
通讯作者:
Miyajima Takaaki
Miyajima Takaaki
中科院分区:
--
文献类型:
--
作者:
Orsztynowicz Manfred;Amano Hideharu;Kubota Kenichi;Miyajima Takaaki

文献摘要

参考文献

被引文献

相似文献

随着半导体制造工艺的缩减,最近机器上的每触发器(B/F)比率变得更低。基于粒子的计算流体动力学(CFD)方法(如移动粒子模拟(MPS))需要比基于模板的CFD方法更高的B/F比。通过利用时间并行性来降低B/F比的技术在CPU和GPU上基于模板的CFD方法中变得流行。本文还介绍了一种将联合收割机时间分块与模板缓冲相结合的技术,该技术适用于FPGA,性能优于CPU和GPU。另一方面,它已被认为是时间的并行性不能利用在基于粒子的CFD方法。这是因为每个桶中的粒子数量,一个覆盖计算域的三维网格,每个时间步都会改变。在本文中,我们提出了一种技术,利用时间并行MPS方法,基于粒子的CFD方法不可压缩流体。其核心思想是MPS中的桶可以被认为是基于模板的CFD中的模板。这是因为在不可压缩流体的情况下,可以根据经验假设桶中的最大颗粒数。据我们所知,这是第一个研究,利用时间并行的粒子为基础的不可压缩流体方法。我们实现了所提出的技术的时间并行度为三。在Intel HLS中的Intel Arria 10 FPGA上进行了优化,并对性能和资源消耗进行了测试。结果表明,优化后的实现与时间并行度为3的实现相比,没有利用CPU上的时间并行实现的2.1倍的加速比。
While the semiconductor manufacturing process is shrinking, the Bytes per Flop (B/F) ratio on recent machines is becoming lower. Particle-based computational fluid dynamics (CFD) methods such as Moving Particle Simulation (MPS) require a higher B/F ratio than that of stencil-based CFD methods. Techniques to reduce the B/F ratio by exploiting temporal parallelism is becoming popular in stencil-based CFD methods on CPU and GPU. It is also reported that a technique to combine temporal blocking with stencil buffer is suitable for FPGA and can outperform CPU and GPU. On the other hand, it has been considered that temporal parallelism cannot be exploited in the particle-based CFD methods. This is because the number of particles in each bucket, a three-dimensional grid covering a computational domain, changes every time-step. In this paper, we propose a technique to exploit temporal parallelism in MPS method, a particle-based CFD method for incompressive fluid. The key idea is that the buckets in MPS can be considered as stencils in stencil-based CFD. This is because the maximum number of particles in a bucket can be assumed empirically in the case of an incompressible fluid. To the best of our knowledge, this is the first research which exploits temporal parallelism in the particle-based incompressible fluid method. We implemented the proposed technique with a degree of temporal parallelism of three. We also optimized it on Intel Arria10 FPGA in Intel HLS, and measured the performance and resource consumption. The result shows that the optimized implementation with a degree of temporal parallelism of three achieved 2.1 times speedup compared with implementation without exploiting temporal parallelism on CPU.
DOI: 10.1109/icis.2016.7550742
发表时间: 2016-08
期刊: 2016 IEEE/ACIS 15th International Conference on Computer and Information Science (ICIS)
影响因子: --
作者:
Hasitha Muthumala Waidyasooriya;M. Hariyama
通讯作者: Hasitha Muthumala Waidyasooriya;M. Hariyama
模板计算的最佳时间阻塞
DOI: 10.1016/j.procs.2015.05.315
发表时间: 2015
期刊: 2013 IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子: --
作者:
T. Muranushi;J. Makino
通讯作者: J. Makino
DOI: --
发表时间: 2012
期刊:
影响因子: --
作者:
S. Koshizuka;M. Oochi; K. Shibata
通讯作者: K. Shibata
DOI: 10.13182/nse96-a24205
发表时间: 1996-07-01
影响因子: 1.2
作者:
Koshizuka, S;Oka, Y
通讯作者: Oka, Y
用于具有恒定内存带宽的模板计算的简单软处理器的可扩展流阵列
DOI: --
发表时间: 2011
期刊: Proc.19th Annual IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM2011)
影响因子: --
作者:
Kentaro Sano;Yoshiaki Hatsuda;Satoru Yamamoto
通讯作者: Satoru Yamamoto