A practical measure of FPGA floating point acceleration for High Performance Computing
A practical measure of FPGA floating point acceleration for High Performance Computing
复制标题
高性能计算 FPGA 浮点加速的实用测量
DOI:
10.1109/asap.2013.6567570
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
D. Strenski
中科院分区:
文献类型:
--
作者:
John D. Cappello;D. Strenski
A key enabler for Field Programmable Gate Arrays (FPGAs) in High Performance Computing (HPC) has been the addition of hard arithmetic cores. These “slices of DSP” dedicated to accelerated number crunching allow FPGAs to deliver more computing muscle, especially for floating point algorithms. This paper compares how an FPGA's performance in a practical HPC application measures up to its theoretical capacity. The implementation of a floating point matrix multiplication algorithm based on a 12x12 MAC (MultiplyAccumulate) array targeting the Xilinx Virtex 7 XT family is described. Several design techniques were used to ensure uninterrupted systolic operation of the array throughout execution, including a novel approach to handling heavily pipelined accumulators, as well as a scheme for overcoming the inherent inefficiencies of DDR3 memory. The result is a sustained "practical" performance range of 144-180 GFLOPS, compared to the target device's "theoretical" range of 257-290 GFLOPS.