Stratix 10 NX Architecture and Applications

Stratix 10 NX Architecture and Applications
复制标题

Stratix 10 NX 架构和应用

DOI:
--
复制
发表时间:
2021
期刊:
Symposium on Field Programmable Gate Arrays
影响因子:
--
通讯作者:
Sergey Gribok
Sergey Gribok
中科院分区:
--
文献类型:
--
作者:
M. Langhammer;Eriko Nurvitadhi;B. Pasca;Sergey Gribok

文献摘要

被引文献

相似文献

人工智能的出现推动了FPGA上高密度低精度算法的采用。这导致了将算术函数和卷积映射到结构上的新方法,以及对嵌入式DSP块的一些更改。FPGA领域之外的技术也在发展,例如为GPU添加了张量结构,以及引入了许多AI ASSP,所有这些都具有比当前FPGA更高的性能和效率。在本文中,我们将介绍Stratix 10 NX设备(NX),这是一种专门针对AI应用空间优化的FPGA变体。除了标准可编程软逻辑结构的计算能力外,新型DSP模块还提供了AI实现中通常使用的低精度乘法器的密集阵列。该模块的架构针对AI中常见的矩阵-矩阵或向量-矩阵乘法进行了调整,其功能旨在有效地处理小型和大型矩阵。基本精度为INT 8和INT 4,沿着具有支持块浮点FP 16和FP 12数字的共享指数支持。所有加法/累加都可以在INT 32或IEEE 754单精度浮点(FP 32)中完成,多个模块可以级联在一起以支持更大的矩阵。我们还将描述可以聚合较小精度乘法器以创建更适用于标准信号处理要求的较大乘法器的方法。在总体计算吞吐量方面,Stratix 10 NX在600 MHz时达到143 INT 8/FP 16 TOPS/FLOP或286 INT 4/FP 12 TOPS/FLOP。根据配置的不同,功率效率在1-4 TOP或TFLOPs/W的范围内。
The advent of AI has driven the adoption of high density low precision arithmetic on FPGAs. This has resulted in new methods in mapping both arithmetic functions as well as dataflows onto the fabric, as well as some changes to the embedded DSP Blocks. Technologies outside of the FPGA realm have also evolved, such as the addition of tensor structures for GPUs, and also the introduction of numerous AI ASSPs, all of which have a higher claimed performance and efficiency than current FPGAs. In this paper we will introduce the Stratix 10 NX device (NX), which is a variant of FPGA specifically optimized for the AI application space. In addition to the computational capabilities of the standard programmable soft logic fabric, a new type of DSP Block provides the dense arrays of low precision multipliers typically used in AI implementations. The architecture of the block is tuned for the common matrix-matrix or vector-matrix multiplications in AI, with capabilities designed to work efficiently for both small and large matrix sizes. The base precisions are INT8 and INT4, along with shared exponent support for support block floating point FP16 and FP12 numerics. All additions/accumulations can be done in INT32 or IEEE754 single precision floating point (FP32), and multiple blocks can be cascaded together to support larger matrices. We will also describe methods by which the smaller precision multipliers can be aggregated to create larger multiplier that are more applicable to standard signal processing requirements. In terms of overall compute throughput, Stratix 10 NX achieves 143 INT8/FP16 TOPs/FLOPs, or 286 INT4/FP12 TOPS/FLOPs at 600MHz. Depending on the configuration, power efficiency is in the range of 1-4 TOPs or TFLOPs/W.