Software-managed automatic data sharing for Coarse-Grained Reconfigurable coprocessors

Software-managed automatic data sharing for Coarse-Grained Reconfigurable coprocessors
复制标题

粗粒度可重构协处理器的软件管理自动数据共享

DOI:
10.1109/fpt.2012.6412148
复制
发表时间:
2012
期刊:
2012 International Conference on Field-Programmable Technology
影响因子:
--
通讯作者:
Jongeun Lee
Jongeun Lee
中科院分区:
--
文献类型:
--
作者:
Toan X. Mai;Jongeun Lee

文献摘要

参考文献

被引文献

相似文献

混合系统中的粗粒度可重构架构 (CGRA) 可以显着加速应用程序计算密集型内核的执行。然而,主处理器(MP)和CGRA之间的数据通信开销可能很大并且会抵消CGRA的加速。在本文中,我们通过使用称为可配置范围内存(CRM)的特殊共享内存提供半自动数据共享技术,解决了减少混合系统中数据通信开销的问题。与之前的工作不同,我们这里使用的 CRM 架构基于比较器,这在阵列可以放置在 CRM 中的位置方面提供了更高的灵活性,同时也使 CRM 的运行时软件管理更具挑战性。我们提出了一种基于首次拟合启发式的高效运行时算法。我们的实验结果表明,与基于 ScratchPad Memory (SPM) 的系统相比,基于 CRM 的系统可以将 MP 和 CGRA 之间的数据传输量减少高达 89.5%,而在仅 MP 执行中,软件管理开销平均仅为内核周期的 1.20~1.34%(取决于 CRM 架构参数)。总体而言,我们基于 CRM 的系统比仅 MP 执行的平均内核加速可提高 3.47 倍,比基于 SPM 的系统提高约 20%。
Coarse-Grained Reconfigurable Architecture (CGRA) in a hybrid system can significantly accelerate the execution of compute-intensive kernels of applications. However, the data communication overhead between the main processor (MP) and the CGRA may be huge and can negate the speed-up of the CGRA. In this paper we address the problem of reducing the data communication overhead in a hybrid system by offering a partially automatic data sharing technique using a special shared memory called Configurable Range Memory (CRM). Unlike the previous work the CRM architecture we use here is based on comparators, which gives much higher flexibility in terms of where an array can be placed within a CRM while it makes the runtime software management of a CRM much more challenging. We present an efficient runtime algorithm based on first-fit heuristic. Our experimental results demonstrate that our CRM-based system can reduce the amount of data transfer between a MP and a CGRA up to 89.5% compared to ScratchPad Memory (SPM)-based systems, while the software management overhead is only 1.20~1.34% on average (depending on CRM architecture parameters) of the kernel cycles in the MP-only execution. Overall our CRM-based system can achieve average kernel speedup of 3.47 times over the MP-only execution, which is about 20% improvement over the SPM-based system.
搜索与 B/R4RGS 结合的支架蛋白
DOI: --
发表时间: 2006
期刊:
影响因子: --
作者:
藤井 聖司;牟田 貴里子;齊藤 修
通讯作者: 齊藤 修