OverGen: Improving FPGA Usability through Domain-specific Overlay Generation

OverGen: Improving FPGA Usability through Domain-specific Overlay Generation
复制标题

DOI:
10.1109/micro56248.2022.00018
复制
发表时间:
2022-10
期刊:
2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Sihao Liu;Jian Weng;Dylan Kupsh;Atefeh Sohrabizadeh;Zhengrong Wang;Licheng Guo;Jiuyang Liu;Maxim Zhulin;Rishabh Mani;Lu Zhang;J. Cong;Tony Nowatzki
Sihao Liu;Jian Weng;Dylan Kupsh;Atefeh Sohrabizadeh;Zhengrong Wang;Licheng Guo;Jiuyang Liu;Maxim Zhulin;Rishabh Mani;Lu Zhang;J. Cong;Tony Nowatzki
中科院分区:
其他
文献类型:
--
作者:
Sihao Liu;Jian Weng;Dylan Kupsh;Atefeh Sohrabizadeh;Zhengrong Wang;Licheng Guo;Jiuyang Liu;Maxim Zhulin;Rishabh Mani;Lu Zhang;J. Cong;Tony Nowatzki

文献摘要

被引文献

相似文献

FPGA已被证明是跨多种类型工作负载的强大计算加速器。主流的编程方法是高级综合(HLS),它将高级语言(例如C+ #pragmas)映射到硬件。不幸的是,HLS在可重配置性、定制性和通用性方面留下了显著的可编程性差距:尽管HLS编译速度很快,但下游物理设计需要数小时到数天; FPGA重配置时间限制了硬件的时间复用能力,工具无法推理跨工作负载的灵活性。覆盖架构通过将可编程设计(例如,CPU、GPU等)映射到可编程设计来减轻上述问题。在FPGA之上。然而,覆盖和FPGA之间的抽象差距导致低效率/利用率。我们的基本思想是开发一个硬件生成框架,针对高度可定制的覆盖,使抽象的差距,可以通过调整设计实例,感兴趣的应用程序降低。我们利用并扩展了先前在可定制空间架构、SoC生成、加速器编译器和设计空间探索器方面的工作,以创建端到端FPGA加速系统。我们的新技术解决了片上存储器和处理元件之间的低效网络,并通过减少所需的重新编译量来改善DSE。我们的框架,OverGen,是非常有竞争力的固定功能的基于HLS的设计,即使生成的设计是可编程的快速重新配置。我们将其与最先进的基于DSE的HLS框架AutoDSE进行了比较。如果不对AutoDSE进行内核调优,OverGen将获得1.2$\times$geoman的性能,即使对基线进行手动内核调优,OverGen仍将获得0.55$\times$geoman的性能--同时提供跨工作负载的运行时灵活性。
FPGAs have been proven to be powerful computational accelerators across many types of workloads. The mainstream programming approach is high level synthesis (HLS), which maps high-level languages (e.g. C+ #pragmas) to hardware. Unfortunately, HLS leaves a significant programmability gap in terms of reconfigurability, customization and versatility: Although HLS compilation is fast, the downstream physical design takes hours to days; FPGA reconfiguration time limits the time-multiplexing ability of hardware, and tools do not reason about cross-workload flexibility. Overlay architectures mitigate the above by mapping a programmable design (e.g. CPU, GPU, etc.) on top of FPGAs. However, the abstraction gap between overlay and FPGA leads to low efficiency/utilization. Our essential idea is to develop a hardware generation framework targeting a highly-customizable overlay, so that the abstraction gap can be lowered by tuning the design instance to applications of interest. We leverage and extend prior work on customizable spatial architectures, SoC generation, accelerator compilers, and design space explorers to create an end-to-end FPGA acceleration system. Our novel techniques address inefficient networks between on-chip memories and processing elements, as well as improving DSE by reducing the amount of recompilation required. Our framework, OverGen, is highly competitive with fixed-function HLS-based designs, even though the generated designs are programmable with fast reconfiguration. We compared to a state-of-the-art DSE-based HLS framework, AutoDSE. Without kernel-tuning for AutoDSE, OverGen gets 1.2$\times$ geomean performance, and even with manual kernel-tuning for the baseline, OverGen still gets 0.55$\times$ geomean performance--all while providing runtime flexibility across workloads.