Mapping for Maximum Performance on FPGA DSP Blocks

Mapping for Maximum Performance on FPGA DSP Blocks
复制标题

映射以实现 FPGA DSP 模块的最高性能

DOI:
10.1109/tcad.2015.2474363
复制
发表时间:
2016
影响因子:
2.9
通讯作者:
Suhaib A. Fahmy
Suhaib A. Fahmy
中科院分区:
计算机科学3区
文献类型:
--
作者:
Bajaj Ronak;Suhaib A. Fahmy

文献摘要

被引文献

相似文献

现代现场可编程门阵列(FPGA)上的数字信号处理(DSP)模块功能强大,支持各种不同的数据路径配置。不幸的是,综合工具中的推理可能无法产生达到最大DSP块吞吐量的电路。我们开发了一个工具,将add/sub/mult节点的图形映射到Xilinx FPGA上的DSP块,以确保最大的吞吐量。这是通过延迟调度直到图形已经被划分到DSP块上并且基于它们的流水线结构被调度之后来完成的,从而导致吞吐量优化的实现。我们的工具准备了各种其他方法的等效实现,包括高级合成(HLS)进行比较。我们表明,所提出的方法提供了一个改进的频率100%的标准流水线代码,和23%的Vivado HLS合成实现,同时保留代码的可移植性,在逻辑资源使用的适度增加的成本。
The digital signal processing (DSP) blocks on modern field programmable gate arrays (FPGAs) are highly capable and support a variety of different datapath configurations. Unfortunately, inference in synthesis tools can fail to result in circuits that reach maximum DSP block throughput. We have developed a tool that maps graphs of add/sub/mult nodes to DSP blocks on Xilinx FPGAs, ensuring maximum throughput. This is done by delaying scheduling until after the graph has been partitioned onto DSP blocks and scheduled based on their pipeline structure, resulting in a throughput optimized implementation. Our tool prepares equivalent implementations in a variety of other methods, including high-level synthesis (HLS) for comparison. We show that the proposed approach offers an improvement in frequency of 100% over standard pipelined code, and 23% over Vivado HLS synthesis implementation, while retaining code portability, at the cost of a modest increase in logic resource usage.