OpenCL-Based FPGA Design to Accelerate the Nodal Discontinuous Galerkin Method for Unstructured Meshes

OpenCL-Based FPGA Design to Accelerate the Nodal Discontinuous Galerkin Method for Unstructured Meshes
复制标题

DOI:
10.1109/fccm.2018.00037
复制
发表时间:
2018-04
期刊:
2018 IEEE 26th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)
影响因子:
--
通讯作者:
Tobias Kenter;Gopinath Mahale;Samer Alhaddad;Y. Grynko;Christian Schmitt;Ayesha Afzal;Frank Hannig;J. Förstner;Christian Plessl
Tobias Kenter;Gopinath Mahale;Samer Alhaddad;Y. Grynko;Christian Schmitt;Ayesha Afzal;Frank Hannig;J. Förstner;Christian Plessl
中科院分区:
其他
文献类型:
--
作者:
Tobias Kenter;Gopinath Mahale;Samer Alhaddad;Y. Grynko;Christian Schmitt;Ayesha Afzal;Frank Hannig;J. Förstner;Christian Plessl

文献摘要

相似文献

迄今为止,FPGA作为科学模拟加速器的探索主要集中在处理规则数据结构的方法的小内核上,例如以有限差分方法的模板计算的形式。在计算科学中,通常采用更先进的方法,以保证更好的稳定性,收敛性,局部性和缩放性。非结构网格被证明是更有效和更准确的,相比,规则的网格,在表示各种形状的计算域。使用非结构化网格,不连续伽辽金方法保留了在时域中进行显式局部更新操作的能力。在这项工作中,我们调查FPGA作为目标平台的节点不连续Galerkin方法的实施,找到时域解的麦克斯韦方程组在非结构化网格。当最大限度地提高数据重用和拟合常数系数到适当分区的片上存储器,高计算强度使我们能够实现和饲料宽数据路径与数百个浮点运算符。通过将片外存储器访问与计算解耦,即使对于部分应用程序所需的不规则访问模式,也可以维持高存储器带宽。使用英特尔/Altera OpenCL SDK的FPGA,我们提出了不同的实施方案的不同多项式阶的方法。在算法的不同阶段,几乎达到了Arria 10平台的计算或带宽限制,因此比高度多线程的CPU实现高出约2倍。
The exploration of FPGAs as accelerators for scientific simulations has so far mostly been focused on small kernels of methods working on regular data structures, for example in the form of stencil computations for finite difference methods. In computational sciences, often more advanced methods are employed that promise better stability, convergence, locality and scaling. Unstructured meshes are shown to be more effective and more accurate, compared to regular grids, in representing computation domains of various shapes. Using unstructured meshes, the discontinuous Galerkin method preserves the ability to perform explicit local update operations for simulations in the time domain. In this work, we investigate FPGAs as target platform for an implementation of the nodal discontinuous Galerkin method to find time-domain solutions of Maxwell's equations in an unstructured mesh. When maximizing data reuse and fitting constant coefficients into suitably partitioned on-chip memory, high computational intensity allows us to implement and feed wide data paths with hundreds of floating point operators. By decoupling off-chip memory accesses from the computations, high memory bandwidth can be sustained, even for the irregular access pattern required by parts of the application. Using the Intel/Altera OpenCL SDK for FPGAs, we present different implementation variants for different polynomial orders of the method. In different phases of the algorithm, either computational or bandwidth limits of the Arria 10 platform are almost reached, thus outperforming a highly multithreaded CPU implementation by around 2x.