The Cost of Flexibility: Embedded versus Discrete Routers in CGRAs for HPC
The Cost of Flexibility: Embedded versus Discrete Routers in CGRAs for HPC
复制标题
灵活性的成本:HPC CGRA 中的嵌入式路由器与离散路由器
DOI:
10.1109/cluster51413.2022.00046
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
and Kentaro Sano
中科院分区:
文献类型:
--
作者:
Boma Adhi;Carlos Cortes;Yiyu Tan;Takuya Kojima;Artur Podobas;and Kentaro Sano
Coarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable architectures that inherit the performance and usability properties of Central Processing Units (CPUs) and the reconfigurability aspects of Field-Programmable Gate Arrays (FPGAs). Historically, CGRAs have been successfully used to accelerate embedded applications and are today also being considered to accelerate High-Performance Computing (HPC) applications in future supercomputers. However, embedded systems and supercomputers are two vastly different domains with different applications and constraints, and it is today not fully understood what CGRA design decisions adequately cater to the HPC market. One such unknown design decision is regarding the interconnect that facilitates intra-CGRA communication. Today, intra-CGRA communication comes in two flavors: using routers closely embedded into the compute units or using discrete routers outside the compute units. The former trades flexibility for a reduction in hardware cost, while the latter has greater flexibility but is more resource hungry. In this paper, we aspire to understand which of both designs best suits the CGRA HPC segment. We extend our previous methodology, which consists of both a parameterized CGRA design and an OpenMPcapable compiler, to accommodate both types of routing designs, including verification tests using RTL simulation. Our results show that the discrete router design can facilitate better use of processing elements (PEs) compared to embedded routers and can achieve up to 79.27% reduction in unnecessary PE occupancy for an aggressively unrolled stencil kernel on a 18 × 16 CGRA at a (estimated) hardware resource overhead cost of 6.3x. This reduction in PE occupancy can be used, for example, to exploit instruction-level parallelism (ILP) through even more aggressive unrolling.
登录
查看更多内容
影响因子:
3.9
作者:
Stephen Friedman;Allan Carroll;B. V. Essen;Benjamin Ylvisaker;C. Ebeling;S. Hauck
通讯作者:
Stephen Friedman;Allan Carroll;B. V. Essen;Benjamin Ylvisaker;C. Ebeling;S. Hauck
DOI:
--
发表时间:
2020
期刊:
IEEE International Conference on Application-Specific Systems, Architectures, and Processors
影响因子:
--
作者:
Artur Podobas;K. Sano;S. Matsuoka
通讯作者:
S. Matsuoka
DOI:
10.1109/ipdps.2019.00019
发表时间:
2018
期刊:
2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
作者:
Jens Domke;Kazuaki Matsumura;M. Wahib;Haoyu Zhang;Keita Yashima;Toshiki Tsuchikawa;Yohei Tsuji;Artur Podobas;S. Matsuoka
通讯作者:
S. Matsuoka
DOI:
--
发表时间:
1999
期刊:
--
影响因子:
--
作者:
A. Agarwal;Saman P. Amarasinghe;R. Barua;M. Frank;W. Lee;Vivek Sarkar;D. Srikrishna;M. Taylor
通讯作者:
M. Taylor
影响因子:
3.9
作者:
Artur Podobas;K. Sano;S. Matsuoka
通讯作者:
Artur Podobas;K. Sano;S. Matsuoka