OpenCL Implementation of Cannon’s Matrix Multiplication Algorithm on Intel Stratix 10 FPGAs

OpenCL Implementation of Cannon’s Matrix Multiplication Algorithm on Intel Stratix 10 FPGAs
复制标题

Cannon 矩阵乘法算法在 Intel Stratix 10 FPGA 上的 OpenCL 实现

DOI:
10.1109/icfpt47387.2019.00020
复制
发表时间:
2019
期刊:
2019 International Conference on Field-Programmable Technology (ICFPT)
影响因子:
--
通讯作者:
C. Plessl
C. Plessl
中科院分区:
--
文献类型:
--
作者:
P. Gorlani;T. Kenter;C. Plessl

文献摘要

参考文献

被引文献

相似文献

Stratix 10 FPGA 卡在加速 HPC 工作负载方面具有巨大潜力,因为 Stratix 10 产品线引入了具有大量 DSP 和内存块的设备。 OpenCL 代码的高级综合可以在 HPC 中的 FPGA 中发挥基础作用,因为与手动优化的 HDL 相比,它允许以更少的开发工作量实现不同的设计。然而,Stratix 10 卡仍然难以使用适用于 OpenCL 的英特尔 FPGA SDK 来充分利用。具有数千个并发算术运算的设计的实现通常会遇到布局和布线问题,这些问题限制了最大频率或完全阻止了成功的综合。为了克服矩阵乘法实现中的这些问题,我们制定了Cannon矩阵乘法算法,并考虑其在FPGA逻辑中的高效综合。我们获得了一个两级块算法,其中较低级别的子矩阵使用我们的 Cannon 算法实现相乘。遵循这种具有多个计算单元的设计方法,我们能够通过 DSP 和内存块的高利用率获得接近或高于 300 MHz 的最大频率。这允许超过 1 TeraFLOPS 的性能结果。
Stratix 10 FPGA cards have a good potential for the acceleration of HPC workloads since the Stratix 10 product line introduces devices with a large number of DSP and memory blocks. The high level synthesis of OpenCL codes can play a fundamental role for FPGAs in HPC, because it allows to implement different designs with lower development effort compared to hand optimized HDL. However, Stratix 10 cards are still hard to fully exploit using the Intel FPGA SDK for OpenCL. The implementation of designs with thousands of concurrent arithmetic operations often suffers from place and route problems that limit the maximum frequency or entirely prevent a successful synthesis. In order to overcome these issues for the implementation of the matrix multiplication, we formulate Cannon's matrix multiplication algorithm with regard to its efficient synthesis within the FPGA logic. We obtain a two-level block algorithm, where the lower level sub-matrices are multiplied using our Cannon's algorithm implementation. Following this design approach with multiple compute units, we are able to get maximum frequencies close to and above 300 MHz with high utilization of DSP and memory blocks. This allows for performance results above 1 TeraFLOPS.
DOI: 10.1109/fpt.2018.00018
发表时间: 2018
期刊: 2018 International Conference on Field-Programmable Technology (FPT)
影响因子: --
作者:
A. Sanaullah;Rushi Patel;M. Herbordt
通讯作者: M. Herbordt
通过重新配置在 HERA 架构上利用混合模式并行性进行矩阵运算
DOI: 10.1049/ip-cdt:20045136
发表时间: 2006
期刊: 22nd International Conference on Field Programmable Logic and Applications (FPL)
影响因子: --
作者:
Xiaofang Wang;Sotirios G. Ziavras
通讯作者: Sotirios G. Ziavras
Spector:OpenCL FPGA 基准测试套件
DOI: 10.1109/fpt.2016.7929519
发表时间: 2016
期刊: 2016 International Conference on Field-Programmable Technology (FPT)
影响因子: --
作者:
Q. Gautier;Alric Althoff;Pingfan Meng;R. Kastner
通讯作者: R. Kastner
基于损失频率、面积和周期的 FPGA 效率分析方法
DOI: 10.1016/j.jpdc.2017.11.012
发表时间: 2018
期刊: J. Parallel Distributed Comput.
影响因子: --
作者:
J. Lemeire;B. Silva;An Braeken;Jan G. Cornelis;A. Touhafi
通讯作者: A. Touhafi