Exploiting Computation Reuse for Stencil Accelerators.

Exploiting Computation Reuse for Stencil Accelerators.
复制标题

DOI:
10.1109/dac18072.2020.9218680
复制
发表时间:
2020-07
期刊:
Proceedings. Design Automation Conference
影响因子:
--
通讯作者:
Cong J
Cong J
中科院分区:
其他
文献类型:
--
作者:
Chi Y;Cong J

文献摘要

被引文献

相似文献

模板内核是在许多应用领域中广泛使用的一种重要的内核类型。多年来,研究人员一直在研究针对不同目标平台的并行化、通信重用和计算重用的优化。然而,由于缺乏完整的设计空间探索和有效的设计空间剪枝,仍然存在挑战,特别是在加速器的计算重用问题上。在这篇文章中,我们给出了一系列模板核(即具有约简操作的模板)的上述挑战的解决方案,其中计算重用模式由于其交换和结合属性而非常灵活。我们形式化地定义了完备的设计空间,并在此基础上给出了一个可证明最优的动态规划算法和一个启发式波束搜索算法,该算法在一个架构感知模型下提供了近最优解。实验结果表明,对于模板核到FPGA的综合,与目前不具备计算重用能力的模板编译器相比,我们提出的算法可以将查找表(LUT)和数字信号处理器(DSP)的使用量分别平均减少58.1%和54.6%,从而使计算密集型内核的平均加速比达到2.3倍,优于最新的CPU/GPU结果。
Stencil kernel is an important type of kernel used extensively in many application domains. Over the years, researchers have been studying the optimizations on parallelization, communication reuse, and computation reuse for various target platforms. However, challenges still exist, especially on the computation reuse problem for accelerators, due to the lack of complete design-space exploration and effective design-space pruning. In this paper, we present solutions to the above challenges for a wide range of stencil kernels (i.e., stencil with reduction operations), where the computation reuse patterns are extremely flexible due to the commutative and associative properties. We formally define the complete design space, based on which we present a provably optimal dynamic programming algorithm and a heuristic beam search algorithm that provides near-optimal solutions under an architecture-aware model. Experimental results show that for synthesizing stencil kernels to FPGAs, compared with state-of-the-art stencil compiler without computation reuse capability, our proposed algorithm can reduce the look-up table (LUT) and digital signal processor (DSP) usage by 58.1% and 54.6% on average respectively, which leads to an average speedup of 2.3× for compute-intensive kernels, outperforming the latest CPU/GPU results.