A practical measure of FPGA floating point acceleration for High Performance Computing

A practical measure of FPGA floating point acceleration for High Performance Computing
复制标题

高性能计算 FPGA 浮点加速的实用测量

DOI:
10.1109/asap.2013.6567570
复制
发表时间:
2013
期刊:
2013 IEEE 24th International Conference on Application-Specific Systems, Architectures and Processors
影响因子:
--
通讯作者:
D. Strenski
D. Strenski
中科院分区:
--
文献类型:
--
作者:
John D. Cappello;D. Strenski

文献摘要

被引文献

相似文献

现场可编程门阵列(FGA)在高性能计算(HPC)中的一个关键推动因素是增加了硬算术核心。这些专门用于加速数字运算的“DSP片”使现场可编程门阵列能够提供更多的计算能力,特别是对于浮点算法。本文比较了在实际高性能计算应用中,现场可编程门阵列的性能如何达到其理论容量。描述了一种基于Xilinx Virtex 7 XT系列的12×12 MAC(MultiplyAcumulate)阵列的浮点矩阵乘法算法的实现。使用了几种设计技术来确保阵列在整个执行过程中不间断的脉动操作,包括一种处理大量流水线累加器的新方法,以及一种克服DDR3存储器固有低效的方案。结果是持续的“实际”性能范围为144-180 GFLOPS,而目标设备的“理论”范围为257-290 GFLOPS。
A key enabler for Field Programmable Gate Arrays (FPGAs) in High Performance Computing (HPC) has been the addition of hard arithmetic cores. These “slices of DSP” dedicated to accelerated number crunching allow FPGAs to deliver more computing muscle, especially for floating point algorithms. This paper compares how an FPGA's performance in a practical HPC application measures up to its theoretical capacity. The implementation of a floating point matrix multiplication algorithm based on a 12x12 MAC (MultiplyAccumulate) array targeting the Xilinx Virtex 7 XT family is described. Several design techniques were used to ensure uninterrupted systolic operation of the array throughout execution, including a novel approach to handling heavily pipelined accumulators, as well as a scheme for overcoming the inherent inefficiencies of DDR3 memory. The result is a sustained "practical" performance range of 144-180 GFLOPS, compared to the target device's "theoretical" range of 257-290 GFLOPS.