FDMAX: An Elastic Accelerator Architecture for Solving Partial Differential Equations

FDMAX: An Elastic Accelerator Architecture for Solving Partial Differential Equations
复制标题

FDMAX:用于求解偏微分方程的弹性加速器架构

DOI:
10.1145/3579371.3589083
复制
发表时间:
2023
期刊:
Proceedings International Symposium on Computer Architecture
影响因子:
--
通讯作者:
Wang, Ke
Wang, Ke
中科院分区:
--
文献类型:
--
作者:
Li, Jiajun;Zhang, Yuxuan;Zheng, Hao;Wang, Ke

文献摘要

参考文献

被引文献

相似文献

偏微分方程在许多科学和工程领域中被广泛用于描述自然现象。许多偏微分方程没有解析解,因此,数值方法已成为流行的近似偏微分方程的解决方案。最广泛使用的数值方法是有限差分法(FDM),它需要精细的网格和高精度的数值迭代,这是计算和存储密集型的。在文献中已经提出了偏微分方程求解加速器,然而,它们通常集中在具有刚性网格尺寸的特定类型的偏微分方程上,这限制了它们更广泛的适用性。此外,他们很少提供洞察到并行计算和数据访问的优化求解偏微分方程,这阻碍了进一步提高性能和能源efficiency.This本文介绍了FDMAX,弹性加速器,有效地支持FDM为不同类型的偏微分方程与任何网格大小。FDMAX采用定制的处理元件(PE)阵列架构,以最小的互连开销最大限度地提高数据重用。PE阵列可以重新配置为分成一组子阵列,以适应不同的网格大小,从而获得最佳效率。此外,PE阵列利用计算和数据重用来提高性能和能量效率,并且可重新配置以支持广泛的PDE,例如椭圆、抛物线和双曲方程。在四个著名的PDE上进行评估,我们的模拟结果显示,FDMAX在Intel Xeon CPU上实现了平均1189倍的加速比,能耗降低了1123倍,在NVIDIA RTX3090 GPU上实现了4.9倍的加速比,能耗降低了6.3倍,在最先进的PDE求解加速器Alrescha上实现了2.9倍的加速比。
Partial Differential Equations (PDEs) are widely employed to describe natural phenomena in many science and engineering fields. Many PDEs do not have analytical solutions, hence, numerical methods have become prevalent for approximating PDE solutions. The most widely used numerical method is the Finite Difference Method (FDM), which requires fine grids and high-precision numerical iterations that are both compute- and memory-intensive. PDE-solving accelerators have been proposed in the literature, however, they usually focus on specific types of PDEs with rigid grid sizes which limits their broader applicability. Besides, they rarely provided insight into the optimizations of parallel computing and data accesses for solving PDEs, which hinders further improvements in performance and energy efficiency.This paper presents FDMAX, an elastic accelerator to efficiently support FDM for different types of PDEs with any grid size. FDMAX employs a customized Processing Element (PE) array architecture that maximizes data reuse with minimized interconnection overhead. The PE array can be reconfigured to break into a set of subarrays to adapt to different grid sizes for optimal efficiency. Moreover, the PE array exploits computation and data reuse for increased performance and energy efficiency, and is reconfigurable to support a wide range of PDEs such as elliptic, parabolic, and hyperbolic equations. Evaluated on four well-known PDEs, our simulation results show that FDMAX achieves on average 1189× speedup with 1123× energy reduction over Intel Xeon CPU, and 4.9× speedup with 6.3× energy reduction over NVIDIA RTX3090 GPU, and 2.9× speedup over Alrescha, the state-of-the-art PDE-solving accelerator.
DOI: 10.1137/0913035
发表时间: 1992-03-01
期刊: SIAM JOURNAL ON SCIENTIFIC AND STATISTICAL COMPUTING
影响因子: --
作者:
VANDERVORST, HA
通讯作者: VANDERVORST, HA
DOI: 10.1109/icis.2016.7550742
发表时间: 2016-08
期刊: 2016 IEEE/ACIS 15th International Conference on Computer and Information Science (ICIS)
影响因子: --
作者:
Hasitha Muthumala Waidyasooriya;M. Hariyama
通讯作者: Hasitha Muthumala Waidyasooriya;M. Hariyama
DOI: 10.1016/j.procs.2010.04.203
发表时间: 2010-05
期刊: --
影响因子: --
作者:
G. Markall;D. Ham;P. Kelly
通讯作者: G. Markall;D. Ham;P. Kelly
使用 GPU 求解微分方程
DOI: --
发表时间: 2009
期刊:
影响因子: --
作者:
加藤大地;沢田篤史;張漢明;野呂昌満;丹治裕一
通讯作者: 丹治裕一
DOI: 10.1201/9781003229100-6
发表时间: 2022-02
期刊: Mastering Git
影响因子: --
作者:
Sufyan bin Uzayr
通讯作者: Sufyan bin Uzayr