Analysis of Blocking and Scheduling for FPGA-Based Floating-Point Matrix Multiplication Analyse du blocage et de l’ordonnancement d’une multiplication matricielle à virgule flottante sur un FPGA

Analysis of Blocking and Scheduling for FPGA-Based Floating-Point Matrix Multiplication Analyse du blocage et de l’ordonnancement d’une multiplication matricielle à virgule flottante sur un FPGA
复制标题

基于 FPGA 的浮点矩阵乘法的分块和调度分析 FPGA 上的虚拟浮点矩阵乘法分析

DOI:
--
复制
发表时间:
2014
期刊:
Canadian journal of electrical and computer engineering
影响因子:
--
通讯作者:
Naraig Manjikian
Naraig Manjikian
中科院分区:
--
文献类型:
--
作者:
Ahmad Khayyat;Naraig Manjikian

文献摘要

被引文献

相似文献

针对基于现场可编程门阵列(现场可编程门阵列)的浮点并行矩阵乘法的设计与实现,考虑了分块与调度问题。为了获得高性能,片上存储器保存在计算分块时重复使用的数据,多个算术单元在每个块内并行执行独立的操作。本文的第一个贡献是详细分析了基于片上存储器使用量和考虑的阻塞和调度计算的方法来表征性能的设计空间。并与前人的工作进行了比较,统一了观点。第二个贡献是Altera Stratix IV EP4SGX530C2 FPGA的灵活高性能实施,具有与外部双倍数据速率同步动态RAM(DDR2 SDRAM)存储器的接口。各种配置选项支持不同目标的优化,所得到的配置已在仿真和硬件中得到验证。对于双精度浮点运算,在160MHZ频率下,运算单元可以实现每秒16G浮点运算(GFLOPS)的性能。
This paper considers blocking and scheduling for the design and implementation of field-programmable gate array (FPGA)-based floating-point parallel matrix multiplication in the presence of a memory hierarchy. For high performance, on-chip memory holds data that are reused when the computation is divided into blocks, and multiple arithmetic units perform independent operations within each block in parallel. The first contribution of this paper is a detailed analysis of the design space to characterize performance based on the amount of on-chip memory used and the approaches considered for blocking and scheduling of the computation. A comparison is also made to prior work with a unified view. The second contribution is a flexible high-performance implementation for the Altera Stratix IV EP4SGX530C2 FPGA with an interface to external double-data-rate synchronous dynamic RAM (DDR2 SDRAM) memory. Various configuration options support optimization of different objectives, and the resulting configurations have been verified in simulation and in hardware. For double-precision floating-point, a performance of 16 giga-floating-point operations per second (GFLOPS) is achievable with 64 arithmetic units at 160 MHz.