MuGRA: A Scalable Multi-Grained Reconfigurable Accelerator Powered by Elastic Neural Network

MuGRA: A Scalable Multi-Grained Reconfigurable Accelerator Powered by Elastic Neural Network
复制标题

MuGRA:由弹性神经网络提供支持的可扩展多粒度可重构加速器

DOI:
10.1109/tcsi.2021.3099034
复制
发表时间:
2022
期刊:
IEEE Transactions on Circuits and Systems I: Regular Papers
影响因子:
--
通讯作者:
Nakashima Yasuhiko
Nakashima Yasuhiko
中科院分区:
--
文献类型:
--
作者:
Kan Yirong;Wu Man;Zhang Renyuan;Nakashima Yasuhiko

文献摘要

相似文献

开发海量核心计算架构,以高速、低成本加速全并行任意计算。所提出的架构可在细粒度(任意函数)、中粒度(灵活的函数特征、精度和操作数数量)和粗粒度(内核组织)方面进行重构。通过在硬件上实现大规模新型二分神经网络(BNN),通过将整个 BNN 划分为任意特定的块而无需冗余来进行重新配置。每块 BNN 近似检索任意函数。通过在软件中重新配置 BNN 拓扑,我们可以轻松调整计算内核的维度,而无需重新布线,并在硬件中实现精度和效率之间的广泛权衡。这样就实现了多粒度可重构加速器(MuGRA)。由于 MuGRA 在所有粒度级别上都很灵活,因此每次验证的各种配置都通过丰富的性能成本矩阵选项进行了演示。从FPGA实现结果来看,与其他传统函数逼近方法相比,我们的方法提供的参数存储要求更少。与相关工作的比较证明,我们的加速器有效地降低了计算延迟,并且精度损失很小。
A massive core computing architecture is developed for accelerating arbitrary calculations in fully parallel with high speed and low cost. The proposed architecture is reconfigurable in fine-grained (arbitrary functions), mid-grained (flexible function feature, accuracy, and number of operands), and coarse-grained (organization of cores). By implementing a large scale of novel bisection neural network (BNN) on hardware, the re-configuration is conducted by partitioning entire BNN into any specific pieces without redundancy. Each piece of BNN retrieves the arbitrary function approximately. By reconfiguring the BNN topology in software, we can easily adjust dimensions of the computing kernel without rewiring, and achieve a wide range of trade-offs between accuracy and efficiency in hardware. In this manner, the multi-grained reconfigurable accelerator (MuGRA) is achieved. Since MuGRA is flexible in all grained levels, various configurations for each validation are demonstrated with rich options of performance-cost matrix. From the FPGA implementation results, compared with other traditional function approximation methods, our method provides fewer parameter storage requirements. The comparison against related works proves that our accelerator effectively reduces the calculation latency with slight accuracy loss.