MITRACA: A Next-Gen Heterogeneous Architecture

MITRACA: A Next-Gen Heterogeneous Architecture
复制标题

MITRACA:下一代异构架构

DOI:
10.1109/mcsoc.2019.00050
复制
发表时间:
2019
期刊:
Embedded Multicore/Many-core Systems-on-Chip
影响因子:
--
通讯作者:
and Boku Taisuke
and Boku Taisuke
中科院分区:
--
文献类型:
--
作者:
Ben Abdelhamid Riadh;Yamaguchi Yoshiki;and Boku Taisuke

文献摘要

相似文献

GPU(图形处理单元)和CPU(中央处理单元)具有足够和适当的性能来计算大规模并行应用,如AI,大数据和材料科学。然而,它们的真实的性能远低于理论上的性能。性能下降的主要原因是它们受到有限的存储器带宽和低效的互连拓扑的影响,这些拓扑没有针对这些类型的应用进行优化。因此,从被称为计算效率的真实的计算性能的观点来看,FPGA(现场可编程门阵列)现在正成为用于具有大规模并行计算的这些类型的应用的有吸引力的芯片。FPGA可以有效地提出优化的通信和桥接不同的计算加速器作为定制的硬件。换句话说,基于FPGA的硬件加速器为高性能和高内存带宽提供了方便的解决方案。然而,一个严重的问题是可用性。例如,使用硬件描述语言的FPGA设计是一项细致的任务,需要专门的技能以及很长的上市时间。覆盖架构将成为解决此问题的合适候选方案,因为它提供了一个软件层,通过抽象结构资源简化了FPGA可编程性。因此,本文提出了一种基于紧密连接的多核CGRA(粗粒度可重构架构)的覆盖架构。它将帮助软件工程师无缝地实现他们的应用程序。我们的最终目标不是当前的细粒度FPGA,而是新的中粒度可编程芯片。如果采用ASIC(专用集成电路)实现,由于工作频率的原因,性能将达到目前FPGA实现的至少十倍以上。在本文中,所提出的覆盖系统提供了一个可编程的接口,虚拟化FPGA资源,让潜在的用户专注于高级软件编程。
GPU (Graphics Processing Unit) and CPU (Central Processing Unit) possess a sufficient and appropriate performance to compute massively parallel applications like AI, Big data, and material sciences. However, their real performance is far lower than those theoretical ones. The primary reason for the performance degradation is that they suffer from limited memory bandwidth and inefficient interconnection topology not optimized for these types of applications. Thus, from the viewpoint of real computational performance called computational efficiency, FPGA (Field Programmable Gate Array) is now becoming an attractive chip for these types of applications with massively parallel computation. FPGA can efficiently propose optimized communication and bridge different computing accelerators as customized hardware. In other words, FPGA-based hardware accelerators offer a convenient solution for both high performance and high memory bandwidth. However, one serious concern is usability. For example, the FPGA design using hardware description language is a meticulous task and requires specialized skill sets as well as a long time to market. An overlay architecture will become an appropriate candidate that can resolve this issue because it offers a software layer that simplifies FPGA programmability by abstracting the fabric resources. Thus, this article proposes an overlay architecture based on a tightly-connected many-core-based CGRA (Coarse-Grained Reconfigurable Architecture). It will help software engineers on seamlessly implementing their applications. Our final goal is not on the current fine-grained FPGAs but new middle-to-course-grained programmable chips. If an ASIC (Application-Specific Integrated Circuit) implementation was adopted, the performance would achieve at least ten times higher compared with the current FPGA implementation because of the working frequency. In this article, the proposed overlay system provides a programmable interface that virtualizes FPGA resources and let prospected users focus on high-level software programming.