Exploring the Versal AI Engines for Accelerating Stencil-based Atmospheric Advection Simulation

Exploring the Versal AI Engines for Accelerating Stencil-based Atmospheric Advection Simulation
复制标题

探索 Versal AI 引擎以加速基于模板的大气平流模拟

DOI:
10.1145/3543622.3573047
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Brown N
Brown N
中科院分区:
--
文献类型:
--
作者:
Brown N

文献摘要

参考文献

被引文献

相似文献

AMD Xilinx的新Versal自适应计算加速平台(ACAP)是一种FPGA架构,将可重配置结构与其他片上强化计算资源相结合。AI引擎就是其中之一,通过以高度矢量化的方式运行,它们提供了重要的原始计算,这对包括HPC模拟在内的一系列工作负载都有潜在的好处。然而,这项技术仍处于早期阶段,尚未被证明可以加速HPC代码,缺乏基准测试和最佳实践。本文介绍了一份经验报告,探索将Piacsek和威廉姆斯(PW)平流方案移植到Versal ACAP上,使用芯片的AI引擎来加速计算。平流是一种基于模板的算法,在大气建模中很常见,包括最初开发该方案的几个气象局代码。使用该算法作为工具,我们探索了构建AI引擎计算内核的最佳方法,以及如何最好地将AI引擎与可编程逻辑连接起来。在VCK5000和Alveo U280以及24核Xeon Platinum Cascade Lake CPU和Nvidia V100 GPU上使用VCK5000与非AI引擎FPGA配置进行性能评估时,我们发现,虽然结构和AI引擎之间的通道数量是一个限制,但通过利用ACAP,我们可以使性能比Alveo U280提高一倍。
AMD Xilinx's new Versal Adaptive Compute Acceleration Platform (ACAP) is an FPGA architecture combining reconfigurable fabric with other on-chip hardened compute resources. AI engines are one of these and, by operating in a highly vectorized manner, they provide significant raw compute that is potentially beneficial for a range of workloads including HPC simulation. However, this technology is still early-on, and as yet unproven for accelerating HPC codes, with a lack of benchmarking and best practice.This paper presents an experience report, exploring porting of the Piacsek and Williams (PW) advection scheme onto the Versal ACAP, using the chip's AI engines to accelerate the compute. A stencil-based algorithm, advection is commonplace in atmospheric modelling, including several Met Office codes who initially developed this scheme. Using this algorithm as a vehicle, we explore optimal approaches for structuring AI engine compute kernels and how best to interface the AI engines with programmable logic. Evaluating performance using a VCK5000 against non-AI engine FPGA configurations on the VCK5000 and Alveo U280, as well as a 24-core Xeon Platinum Cascade Lake CPU and Nvidia V100 GPU, we found that whilst the number of channels between the fabric and AI engines are a limitation, by leveraging the ACAP we can double performance compared to an Alveo U280.
在 Xilinx 和 Intel FPGA 上加速大气建模的平流
DOI: 10.1109/cluster48925.2021.00113
发表时间: 2021
期刊: --
影响因子: --
作者:
Brown N
通讯作者: Brown N
Xilinx Versal ACAP 重离子辐照的初步结果。
DOI: 10.2172/1884175
发表时间: 2021
期刊: Proposed for presentation at the Single Event Effects Symposium and Military and Aerospace Programmable Logic Devices Workshop held August 31-September 2, 2021 in Virtual, Virtual.
影响因子: --
作者:
David Lee;Gregory Allen;M. Cannon;Hunter Earnest;Paul Thelen;Nathaniel Dodds;Jeff McCasland;Carol Chen
通讯作者: Carol Chen