A 1.15 TOPS/W, 16-Cores Parallel Ultra-Low Power Cluster with 2b-to-32b Fully Flexible Bit-Precision and Vector Lockstep Execution Mode

A 1.15 TOPS/W, 16-Cores Parallel Ultra-Low Power Cluster with 2b-to-32b Fully Flexible Bit-Precision and Vector Lockstep Execution Mode
复制标题

1.15 TOPS/W、16 核并行超低功耗集群,具有 2b 至 32b 完全灵活的位精度和矢量锁步执行模式

DOI:
--
复制
发表时间:
2021
期刊:
European Solid-State Circuits Conference
影响因子:
--
通讯作者:
D. Rossi
D. Rossi
中科院分区:
--
文献类型:
--
作者:
Angelo Garofalo;G. Ottavi;Alfio Di Mauro;Francesco Conti;Giuseppe Tagliavini;L. Benini;D. Rossi

文献摘要

被引文献

相似文献

物联网终端节点需要极高的性能和能源效率,再加上高度的灵活性,以应对日益增长的计算需求和各种现代近传感器数据分析应用。在线性代数、深度神经网络(DNN)推理和在线学习等多个领域,低位宽和混合精度算法正成为解决近传感器分析挑战的一种趋势。我们介绍了Dustin,一个完全可编程的多指令多数据(MIMD)集群,集成了16个RISC-V内核,具有2b到32b位精度指令集架构(ISA)扩展,可实现细粒度可调混合精度计算,比最先进的完全可编程设备提高3.7倍和1.9倍的性能和效率。集群可以动态配置为矢量锁步执行模式(VLEM),关闭除一个之外的所有中频阶段,在不降低性能的情况下降低功耗高达38%。该集群采用65nm CMOS技术实现,峰值性能为58 GOPS,峰值效率为1.15 TOPS/W。
IoT end-nodes require extreme performance and energy efficiency coupled with high flexibility to deal with the increasing computational requirements and variety of modern near-sensor data analytics applications. Low-Bitwidth and Mixed-Precision arithmetic is emerging as a trend to address the near-sensor analytics challenge in several fields such as linear algebra, Deep Neural Networks (DNN) inference, and on-line learning. We present Dustin, a fully programmable Multiple Instruction Multiple Data (MIMD) cluster integrating 16 RISC-V cores featuring 2b-to-32b bit-precision instruction set architecture (ISA) extensions enabling fine-grain tunable mixed-precision computation, improving performance and efficiency by 3.7 x and 1.9 x over state-of-the-art fully programmable devices. The cluster can be dynamically configured in Vector Lockstep Execution Mode (VLEM), turning off all IF stages except one, reducing power consumption by up to 38% with no performance degradation. The cluster, implemented in 65nm CMOS technology, achieves a peak performance of 58 GOPS and a peak efficiency of 1.15 TOPS/W.