Inter-tile reuse optimization applied to bandwidth constrained embedded accelerators

Inter-tile reuse optimization applied to bandwidth constrained embedded accelerators
复制标题

DOI:
10.7873/date.2015.1033
复制
发表时间:
2015-03
期刊:
2015 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子:
--
通讯作者:
Maurice Peemen;B. Mesman;H. Corporaal
Maurice Peemen;B. Mesman;H. Corporaal
中科院分区:
其他
文献类型:
--
作者:
Maurice Peemen;B. Mesman;H. Corporaal

文献摘要

被引文献

相似文献

采用高级合成(HLS)工具的加速器设计时间大大减少了。剩下的一个复杂的缩放问题是数据传输瓶颈。为了扩展性能加速器,需要大量数据,并且通常受到互连资源的限制。此外,加速器所花费的能量通常由数据传输(以内存参考的形式或互连上的数据移动形式)主导。在本文中,我们通过探索计算重新排序和局部缓冲用法来大大减少加速器通信。因此,我们提出了一种新的分析方法,以优化嵌套循环,以通过循环转换(例如交换和瓷砖)进行嵌入式数据重复使用。我们专注于嵌入式加速器,这些加速器可用于芯片(SOC)的多加速器系统,因此性能,区域和能量是此探索的关键。 1)在图像/视频处理域中的三个常见嵌入式应用程序(Demosaicing,块匹配,对象检测)中,我们表明我们的方法论将数据移动降低到2.1倍,与最佳的瓷砖内部优化相比。 2)我们证明,我们的小加速器(1-3%的FPGA资源)可以将简单的微封面软核提升到高端Intel-I7处理器的性能水平。
The adoption of High-Level Synthesis (HLS) tools has significantly reduced accelerator design time. A complex scaling problem that remains is the data transfer bottleneck. To scale-up performance accelerators require huge amounts of data, and are often limited by interconnect resources. In addition, the energy spent by the accelerator is often dominated by the transfer of data, either in the form of memory references or data movement on interconnect. In this paper we drastically reduce accelerator communication by exploration of computation reordering and local buffer usage. Consequently, we present a new analytical methodology to optimize nested loops for inter-tile data reuse with loop transformations like interchange and tiling. We focus on embedded accelerators that can be used in a multi-accelerator System on Chip (SoC), so performance, area, and energy are key in this exploration. 1) On three common embedded applications in the image/video processing domain (demosaicing, block matching, object detection), we show that our methodology reduces data movement up to 2.1x compared to the best case of intra-tile optimization. 2) We demonstrate that our small accelerators (1-3% FPGA resources) can boost a simple MicroBlaze soft-core to the performance level of a high-end Intel-i7 processor.