The Cost of Flexibility: Embedded versus Discrete Routers in CGRAs for HPC

The Cost of Flexibility: Embedded versus Discrete Routers in CGRAs for HPC
复制标题

灵活性的成本:HPC CGRA 中的嵌入式路由器与离散路由器

DOI:
10.1109/cluster51413.2022.00046
复制
发表时间:
2022
期刊:
Proceedings of IEEE Cluster Conference (CLUSTER)
影响因子:
--
通讯作者:
and Kentaro Sano
and Kentaro Sano
中科院分区:
--
文献类型:
--
作者:
Boma Adhi;Carlos Cortes;Yiyu Tan;Takuya Kojima;Artur Podobas;and Kentaro Sano

文献摘要

参考文献

被引文献

相似文献

粗粒度可重构阵列(CGRA)是一类可重构架构,其继承了中央处理单元(CPU)的性能和可用性属性以及现场可编程门阵列(FPGA)的可重构性方面。从历史上看,CGRA已成功用于加速嵌入式应用程序,如今也被考虑用于加速未来超级计算机中的高性能计算(HPC)应用程序。然而,嵌入式系统和超级计算机是两个截然不同的领域,具有不同的应用和限制,目前尚不完全清楚CGRA的哪些设计决策能够充分满足HPC市场的需求。一个这样的未知设计决策是关于促进CGRA内通信的互连。今天,CGRA内部通信有两种方式:使用紧密嵌入计算单元的路由器或使用计算单元外部的离散路由器。前者以灵活性换取硬件成本的降低,而后者具有更大的灵活性,但更需要资源。在本文中,我们希望了解这两种设计中哪一种最适合CGRA HPC细分市场。我们扩展我们以前的方法,它包括一个参数化的CGRA设计和一个OpenMPable编译器,以适应这两种类型的路由设计,包括使用RTL仿真验证测试。我们的研究结果表明,与嵌入式路由器相比,离散路由器设计可以更好地使用处理元件(PE),并且可以在18 × 16 CGRA上以6.3倍的(估计)硬件资源开销成本将积极展开的模板内核的不必要的PE占用减少高达79.27%。PE占用率的这种减少可以用于例如通过甚至更积极的展开来利用并行级并行(ILP)。
Coarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable architectures that inherit the performance and usability properties of Central Processing Units (CPUs) and the reconfigurability aspects of Field-Programmable Gate Arrays (FPGAs). Historically, CGRAs have been successfully used to accelerate embedded applications and are today also being considered to accelerate High-Performance Computing (HPC) applications in future supercomputers. However, embedded systems and supercomputers are two vastly different domains with different applications and constraints, and it is today not fully understood what CGRA design decisions adequately cater to the HPC market. One such unknown design decision is regarding the interconnect that facilitates intra-CGRA communication. Today, intra-CGRA communication comes in two flavors: using routers closely embedded into the compute units or using discrete routers outside the compute units. The former trades flexibility for a reduction in hardware cost, while the latter has greater flexibility but is more resource hungry. In this paper, we aspire to understand which of both designs best suits the CGRA HPC segment. We extend our previous methodology, which consists of both a parameterized CGRA design and an OpenMPcapable compiler, to accommodate both types of routing designs, including verification tests using RTL simulation. Our results show that the discrete router design can facilitate better use of processing elements (PEs) compared to embedded routers and can achieve up to 79.27% reduction in unnecessary PE occupancy for an aggressively unrolled stencil kernel on a 18 × 16 CGRA at a (estimated) hardware resource overhead cost of 6.3x. This reduction in PE occupancy can be used, for example, to exploit instruction-level parallelism (ILP) through even more aggressive unrolling.
DOI: 10.1145/1508128.1508158
发表时间: 2009-02
期刊: RSC Advances
影响因子: 3.9
作者:
Stephen Friedman;Allan Carroll;B. V. Essen;Benjamin Ylvisaker;C. Ebeling;S. Hauck
通讯作者: Stephen Friedman;Allan Carroll;B. V. Essen;Benjamin Ylvisaker;C. Ebeling;S. Hauck
用于探索粗粒度可重构架构的基于模板的框架
DOI: --
发表时间: 2020
期刊: IEEE International Conference on Application-Specific Systems, Architectures, and Processors
影响因子: --
作者:
Artur Podobas;K. Sano;S. Matsuoka
通讯作者: S. Matsuoka
高性能计算中的双精度 FPU:财富的尴尬?
DOI: 10.1109/ipdps.2019.00019
发表时间: 2018
期刊: 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子: --
作者:
Jens Domke;Kazuaki Matsumura;M. Wahib;Haoyu Zhang;Keita Yashima;Toshiki Tsuchikawa;Yohei Tsuji;Artur Podobas;S. Matsuoka
通讯作者: S. Matsuoka
原始编译器项目
DOI: --
发表时间: 1999
期刊: --
影响因子: --
作者:
A. Agarwal;Saman P. Amarasinghe;R. Barua;M. Frank;W. Lee;Vivek Sarkar;D. Srikrishna;M. Taylor
通讯作者: M. Taylor
DOI: 10.1109/access.2020.3012084
发表时间: 2020-04
期刊: IEEE Access
影响因子: 3.9
作者:
Artur Podobas;K. Sano;S. Matsuoka
通讯作者: Artur Podobas;K. Sano;S. Matsuoka