Parallel and Scalable Custom Computing for Real-Time Fluid Simulation on a Cluster Node with Four Tightly-Coupled FPGAs

Parallel and Scalable Custom Computing for Real-Time Fluid Simulation on a Cluster Node with Four Tightly-Coupled FPGAs
复制标题

在具有四个紧耦合 FPGA 的集群节点上进行实时流体模拟的并行且可扩展的自定义计算

DOI:
10.1109/fpl.2013.6645625
复制
发表时间:
2013
期刊:
Proceedings of the 23rd International Conference on Field-Programmable Logic and Applications (FPL2013)
影响因子:
--
通讯作者:
Hayato Suzuki and Yoshiaki Kono
Hayato Suzuki and Yoshiaki Kono
中科院分区:
--
文献类型:
--
作者:
Kentaro Sano;Ryo Ito;Hayato Suzuki and Yoshiaki Kono

文献摘要

相似文献

仅给出摘要表格。基于计算流体动力学(CFD)的数值模拟现在是一种不可或缺的技术,尤其是在工业领域,因为它能够以比使用风洞进行实验更低的成本获取各种数据。格子玻尔兹曼法(LBM)是CFD方案之一,用于计算包括多相流在内的各种问题。 LBM具有良好的并行性,但同时需要大量数据来计算每个格点,导致运算强度较低。因此,当使用通用处理器和 GPU 进行计算时,LBM 的持续性能受到内存带宽的限制,而不是算术性能的限制。更糟糕的是,互连网络的带宽不足和高延迟导致并行计算的开销较大,尤其是在强扩展的情况下。
Summary form only given. Numerical simulation based on computational fluid dynamics (CFD) is now an indispensable technique especially in industry due to its acquisition capability of various data at a lower cost than experiments using a wind tunnel. The lattice Boltzmann method (LBM) is one of the CFD schemes, which is used to compute various problems including multiphase flow. LBM has good parallelism, but simultaneously requires many data to compute each lattice point, resulting in a low operational intensity. Consequently, the sustained performance of LBM is limited by memory bandwidth rather than arithmetic performance when computed by using general-purpose processors and GPUs. To make matters worse, insufficient bandwidth and high-latency of an interconnection network cause a relatively big overhead in parallel computing, especially in the case of strong-scaling.