PathSeeker: A Fast Mapping Algorithm for CGRAs

PathSeeker: A Fast Mapping Algorithm for CGRAs
复制标题

DOI:
10.23919/date54114.2022.9774520
复制
发表时间:
2022-03
期刊:
2022 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子:
--
通讯作者:
M. Balasubramanian;Aviral Shrivastava
M. Balasubramanian;Aviral Shrivastava
中科院分区:
其他
文献类型:
--
作者:
M. Balasubramanian;Aviral Shrivastava

文献摘要

被引文献

相似文献

由于粗粒度可重构阵列(CGRA)编译器将计算密集型循环高效地映射到2-D阵列上,多年来已成为一种低功耗加速器。当遇到给定节点的映射失败时,现有的映射技术要么退出并重新尝试映射,要么执行回溯,即递归地移除先前映射的节点以找到有效的映射。放弃映射并重新开始可能会降低映射的质量和编译时间。甚至回溯也可能不是最佳选择,因为前一个节点可能不是错误放置的节点。为了解决这个问题,我们提出了Path Seeker-一种映射方法,它分析映射失败并对调度进行局部调整以获得映射。在来自MiBch、Rodinia和Parboil基准测试套件的35个顶级性能关键循环上的实验结果表明,与之前无法分别映射20个和5个循环的最先进方法GraphMinor和Ramp相比,PathSeeker可以更好地映射所有这些循环,并显著减少编译时间。在这些基准测试中,Path Seeker在编译加速比为GraphMinor的550倍时获得了28%的性能提升,在4x4 CGRA上实现了10倍的编译加速比提升了3%的性能。
Coarse-grained reconfigurable arrays (CGRAs) have gained traction over the years as a low-power accelerator due to the efficient mapping of the compute-intensive loops onto the 2-D array by the CGRA compiler. When encountering a mapping failure for a given node, existing mapping techniques either exit and retry the mapping anew, or perform backtracking, i.e., recursively remove the previously mapped node to find a valid mapping. Abandoning mapping and starting afresh can deteriorate the quality of mapping and the compilation time. Even backtracking may not be the best choice since the previous node may not be the incorrectly placed node. To tackle this issue, we propose PathSeeker - a mapping approach that analyzes mapping failures and performs local adjustments to the schedule to obtain a mapping. Experimental results on 35 top performance-critical loops from MiBench, Rodinia, and Parboil benchmark suites demonstrate that PathSeeker can map all of them with better mapping quality and dramatically less compilation time than the previous state-of-the-art approaches - GraphMinor and RAMP, which were unable to map 20 and 5 loops, respectively. Over these benchmarks, PathSeeker achieves 28% better performance at 550x compilation speedup over GraphMinor and 3% better performance at 10x compilation speedup over RAMP on a 4x4 CGRA.