RAMP: Resource-Aware Mapping for CGRAs

RAMP: Resource-Aware Mapping for CGRAs
复制标题

DOI:
10.1145/3195970.3196101
复制
发表时间:
2018-06
期刊:
2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC)
影响因子:
--
通讯作者:
Shail Dave;M. Balasubramanian;Aviral Shrivastava
Shail Dave;M. Balasubramanian;Aviral Shrivastava
中科院分区:
其他
文献类型:
--
作者:
Shail Dave;M. Balasubramanian;Aviral Shrivastava

文献摘要

被引文献

相似文献

粗粒度可重构阵列(CGRA)是一种很有前景的解决方案,甚至可以加速非并行循环。通过 CGRA 实现的加速关键取决于映射(循环操作到 CGRA 的 PE 上)的良好程度,特别是编译器在操作之间路由依赖关系的能力。之前的工作已经探索了几种路由数据依赖关系的机制,包括通过其他 PE、寄存器、内存甚至重新计算进行路由。所有这些路由选项都会更改要映射到 PE 上的图(通常通过添加新操作),并且如果不重新调度,可能无法映射新图。然而,现有技术在编译过程的布局布线 (P&R) 阶段探索这些布线选项,该阶段在调度步骤之后执行。结果,他们要么无法实现映射,要么获得较差的结果。我们的方法 RAMP 在调度步骤之前明确且智能地探索各种路由选项,并提高了映射能力和映射质量。通过评估超过 12 种架构配置的 MiBench 基准测试的顶级性能关键循环,我们发现 RAMP 能够比顺序执行将循环加速 23 倍,与最先进的技术相比,实现了 2.13 倍的几何平均加速。
Coarse-grained reconfigurable array (CGRA) is a promising solution that can accelerate even non-parallel loops. Acceleration achieved through CGRAs critically depends on the goodness of mapping (of loop operations onto the PEs of CGRA), and in particular, the compiler’s ability to route the dependencies among operations. Previous works have explored several mechanisms to route data dependencies, including, routing through other PEs, registers, memory, and even re-computation. All these routing options change the graph to be mapped onto PEs (often by adding new operations), and without re-scheduling, it may be impossible to map the new graph. However, existing techniques explore these routing options inside the Place and Route (P&R) phase of the compilation process, which is performed after the scheduling step. As a result, they either may not achieve the mapping or obtain poor results. Our method RAMP, explicitly and intelligently explores the various routing options, before the scheduling step, and makes improve the mapping-ability and mapping quality. Evaluating top performance-critical loops of MiBench benchmarks over 12 architectural configurations, we find that RAMP is able to accelerate loops by 23× over sequential execution, achieving a geomean speedup of 2.13× over state-of-the-art.