ISOSceles: Accelerating Sparse CNNs through Inter-Layer Pipelining

ISOSceles: Accelerating Sparse CNNs through Inter-Layer Pipelining
复制标题

DOI:
10.1109/hpca56546.2023.10071080
复制
发表时间:
2023-02
期刊:
2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Yifan Yang;J. Emer;Daniel S. Sanchez
Yifan Yang;J. Emer;Daniel S. Sanchez
中科院分区:
其他
文献类型:
--
作者:
Yifan Yang;J. Emer;Daniel S. Sanchez

文献摘要

相似文献

稀疏CNN大大降低了密集CNN的计算和存储成本。但稀疏性也使CNN的数据密集度更高,因为每个值的重用次数更少。因此,当前稀疏CNN加速器每次处理一层,会受到内存流量的影响。我们提出了ISOSceles,这是一种新的稀疏CNN加速器,通过层间流水线大大减少了数据移动:重叠连续层的执行,以便层的输出激活被下一层快速消耗,而不会溢出到芯片外。流水线极大地提高了重用性,但用现有的方法实现是具有挑战性的,这些方法仅限于密集的CNN。ISOSceles依赖于一种新的输入-稳定输出-稳定(IS-OS)并行流,它以相同的顺序消耗输入和产生输出,大大减少了现有并行流的中间大小。ISOSceles有效地实现了IS-OS,并利用时间复用和动态调度来实现多个层的流水线,尽管稀疏性会导致工作的巨大变化。在广泛的稀疏CNN上,ISOSceles的性能比最先进的加速器高出gmean 4.3倍(高达6.7倍),并在使用更少面积的同时减少了4.7倍(高达8.5倍)的流量。
Sparse CNNs dramatically reduce computation and storage costs over dense ones. But sparsity also makes CNNs more data-intensive, as each value is reused fewer times. Thus, current sparse CNN accelerators, which process one layer at a time, are bottlenecked by memory traffic.We present ISOSceles, a new sparse CNN accelerator that dramatically reduces data movement through inter-layer pipelining: overlapping the execution of consecutive layers so that a layer’s output activations are quickly consumed by the next layer without spilling them off-chip. Pipelining greatly increases reuse, but it is challenging to implement with existing approaches, which are limited to dense CNNs. ISOSceles relies on a novel input-stationary output-stationary (IS-OS) dataflow that consumes inputs and produces outputs in the same order, greatly reducing intermediate sizes over existing dataflows. ISOSceles implements IS-OS efficiently and leverages time-multiplexing and dynamic scheduling to pipeline multiple layers despite the large variations in work that sparsity induces.On a wide range of sparse CNNs, ISOSceles outperforms a state-of-the-art accelerator by gmean 4.3× (up to 6.7×), and reduces traffic by 4.7× (up to 8.5×) while using less area.