Porting incompressible flow matrix assembly to FPGAs for accelerating HPC engineering simulations

Porting incompressible flow matrix assembly to FPGAs for accelerating HPC engineering simulations
复制标题

将不可压缩流动矩阵组件移植到 FPGA 以加速 HPC 工程仿真

DOI:
--
复制
发表时间:
2021
期刊:
International Workshop on Heterogeneous High-performance Reconfigurable Computing
影响因子:
--
通讯作者:
Nick Brown
Nick Brown
中科院分区:
--
文献类型:
--
作者:
Nick Brown

文献摘要

参考文献

被引文献

相似文献

工程是超级计算的一个重要领域,Alya模型是进行此类模拟的流行代码。随着越来越多的需求,从用户模型更大,更复杂的系统在减少时间的解决方案,重要的是要探索的作用,新的硬件技术,如FPGA,可以在加速这些工作负载在未来的exascale systems.In本文中,我们探讨了移植的Alya的不可压缩流矩阵组装内核,占很大比例的模型运行时,到FPGA。在详细描述了成功的策略,在内核级的优化,然后我们探讨FPGA和主机CPU之间的工作负载,映射这些技术之间的内核最合适的部分,使我们能够更有效地利用FPGA之间的共享。然后,我们将我们的方法在Xilinx Alveo U280上的性能与24核Xeon Platinum CPU和Nvidia V100 GPU进行了比较,FPGA的性能明显优于CPU,并对GPU进行了性能测试,同时功耗大大降低。这项工作的结果是一份经验报告,描述了适当的低功耗优化,我们认为可以更广泛地应用于HPC代码的案例研究,以及这种特定工作负载的性能比较,证明了FPGA在加速HPC工程模拟方面的潜力。
Engineering is an important domain for supercomputing, with the Alya model being a popular code for undertaking such simulations. With ever increasing demand from users to model larger, more complex systems at reduced time to solution it is important to explore the role that novel hardware technologies, such as FPGAs, can play in accelerating these workloads on future exascale systems.In this paper we explore the porting of Alya’s incompressible flow matrix assembly kernel, which accounts for a large proportion of the model runtime, onto FPGAs. After describing in detail successful strategies for optimisation at the kernel level, we then explore sharing the workload between the FPGA and host CPU, mapping most appropriate parts of the kernel between these technologies, enabling us to more effectively exploit the FPGA. We then compare the performance of our approach on a Xilinx Alveo U280 against a 24-core Xeon Platinum CPU and Nvidia V100 GPU, with the FPGA significantly out-performing the CPU and performing comparably against the GPU, whilst drawing substantially less power. The result of this work is both an experience report describing appropriate dataflow optimisations which we believe can be applied more widely as a case-study across HPC codes, and a performance comparison for this specific workload that demonstrates the potential for FPGAs in accelerating HPC engineering simulations.
将 LFRic 天气和气候模型移植到 EuroExa 架构 FPGA 的第一步
DOI: 10.1155/2019/7807860
发表时间: 2019
影响因子: --
作者:
Ashworth M
通讯作者: Ashworth M