CGRA express: accelerating execution using dynamic operation fusion

CGRA express: accelerating execution using dynamic operation fusion
复制标题

CGRA Express:使用动态操作融合加速执行

DOI:
--
复制
发表时间:
2009
期刊:
International Conference on Compilers, Architecture, and Synthesis for Embedded Systems
影响因子:
--
通讯作者:
S. Mahlke
S. Mahlke
中科院分区:
--
文献类型:
--
作者:
Yongjun Park;Hyunchul Park;S. Mahlke

文献摘要

被引文献

相似文献

粗粒度可重构体系结构(CGRA)通过提供具有高计算吞吐量、可扩展性、低成本和能量效率的潜力的可编程性而呈现出吸引人的硬件平台。CGRA已有效地用于包含大量并行级并行的最内层循环。相反,非循环和外循环代码是延迟受限的,并且不提供大量的并行级并行。在这些情况下,CGRA是无效的,因为大多数资源仍然闲置。在本文中,动态操作融合,使CGRA有效地加速延迟受限的代码区域。通过在常规CGRA中的功能单元之间添加的小旁路网络和子周期模调度器的组合来实现动态操作融合,以自动识别融合的机会。结果表明,在4x4 CGRA上,动态操作融合将总应用程序运行时间减少了17%。
Coarse-grained reconfigurable architectures (CGRAs) present an appealing hardware platform by providing programmability with the potential for high computation throughput, scalability, low cost, and energy efficiency. CGRAs have been effectively used for innermost loops that contain an abundant of instruction-level parallelism. Conversely, non-loop and outer-loop code are latency constrained and do not offer significant amounts of instruction-level parallelism. In these situations, CGRAs are ineffective as the majority of the resources remain idle. In this paper, dynamic operation fusion is introduced to enable CGRAs to effectively accelerate latency-constrained code regions. Dynamic operation fusion is enabled through the combination of a small bypass network added between function units in a conventional CGRA and a sub-cycle modulo scheduler to automatically identify opportunities for fusion. Results show that dynamic operation fusion reduced total application run-time by up to 17% on a 4x4 CGRA.