AutoBridge: Coupling Coarse-Grained Floorplanning and Pipelining for High-Frequency HLS Design on Multi-Die FPGAs.

AutoBridge: Coupling Coarse-Grained Floorplanning and Pipelining for High-Frequency HLS Design on Multi-Die FPGAs.
复制标题

AutoBridge:耦合粗粒度布局和流水线技术,用于多芯片FGA的高频HLS设计。

DOI:
10.1145/3431920.3439289
复制
发表时间:
2021-03
期刊:
FPGA. ACM International Symposium on Field-Programmable Gate Arrays
影响因子:
--
通讯作者:
Cong J
Cong J
中科院分区:
其他
文献类型:
--
作者:
Guo L;Chi Y;Wang J;Lau J;Qiao W;Ustun E;Zhang Z;Cong J

文献摘要

被引文献

相似文献

尽管越来越多地采用高层次综合(HLS)的设计生产力的优势,仍然有一个显着的差距,可实现的频率之间的HLS设计和手工RTL的。限制HLS输出的时序质量的一个关键因素是难以准确估计HLS级的互连延迟。当大型HLS设计在最新的多管芯FPGA上实现时,这个问题变得更加严重。为了应对这一挑战,我们提出了AutoBridge,这是一个自动化框架,它在HLS编译过程中将粗粒度的布局规划步骤与流水线相结合。首先,我们的方法为HLS提供了设计的全局物理布局视图,使HLS能够更容易地识别和流水线的长导线,特别是那些跨越芯片边界。其次,通过利用HLS流水线的灵活性,布图规划器能够在FPGA器件上的多个管芯上分布设计逻辑,而不会降低时钟频率。这防止了放置器将逻辑积极地打包在单个管芯上,这通常导致最终降低时序的局部布线拥塞。由于流水线可能会引入额外的延迟,我们进一步提出的分析和算法,以确保增加的延迟不会损害整体吞吐量。AutoBridge可以集成到Xilinx FPGA的现有CAD工具流中。在我们的实验中,共有43个设计配置,我们提高了平均频率从147 MHz到297 MHz(102%的改善),没有损失的吞吐量和资源利用率的变化可以忽略不计。值得注意的是,在16个实验中,我们使最初不可路由的设计平均达到274 MHz。该工具可在https://github.com/Licheng-Guo/AutoBridge上获得。
Despite an increasing adoption of high-level synthesis (HLS) for its design productivity advantages, there remains a significant gap in the achievable frequency between an HLS design and a handcrafted RTL one. A key factor that limits the timing quality of the HLS outputs is the difficulty in accurately estimating the interconnect delay at the HLS level. This problem becomes even worse when large HLS designs are implemented on the latest multi-die FPGAs. To tackle this challenge, we propose AutoBridge, an automated framework that couples a coarse-grained floorplanning step with pipelining during HLS compilation. First, our approach provides HLS with a view on the global physical layout of the design, allowing HLS to more easily identify and pipeline the long wires, especially those crossing the die boundaries. Second, by exploiting the flexibility of HLS pipelining, the floorplanner is able to distribute the design logic across multiple dies on the FPGA device without degrading clock frequency. This prevents the placer from aggressively packing the logic on a single die which often results in local routing congestion that eventually degrades timing. Since pipelining may introduce additional latency, we further present analysis and algorithms to ensure the added latency will not compromise the overall throughput. AutoBridge can be integrated into the existing CAD toolflow for Xilinx FPGAs. In our experiments with a total of 43 design configurations, we improve the average frequency from 147 MHz to 297 MHz (a 102% improvement) with no loss of throughput and a negligible change in resource utilization. Notably, in 16 experiments we make the originally unroutable designs achieve 274 MHz on average. The tool is available at https://github.com/Licheng-Guo/AutoBridge.