Comprehensive Accelerator-Dataflow Co-design Optimization for Convolutional Neural Networks

Comprehensive Accelerator-Dataflow Co-design Optimization for Convolutional Neural Networks
复制标题

DOI:
10.1109/cgo53902.2022.9741281
复制
发表时间:
2022-04
期刊:
2022 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
影响因子:
--
通讯作者:
Miheer Vaidya;Aravind Sukumaran-Rajam;A. Rountev;P. Sadayappan
Miheer Vaidya;Aravind Sukumaran-Rajam;A. Rountev;P. Sadayappan
中科院分区:
其他
文献类型:
--
作者:
Miheer Vaidya;Aravind Sukumaran-Rajam;A. Rountev;P. Sadayappan

文献摘要

相似文献

用于将卷积神经网络层映射到空间加速器阵列(被称为空间加速器阵列)的可能调度的设计空间是巨大的。关键架构参数(例如处理元件的数量、寄存器文件和暂存存储器的大小)的协同设计沿着用于优化一个或多个CNN级的实现的并行设计使得设计空间爆炸性地更大。最近的几项努力通过启发式或有限的搜索策略解决了CNN加速器的设计空间探索问题。在本文中,我们开发了第一个优化方法,使用分析建模和约束非线性优化问题的解决方案,全面的算法架构协同设计优化。使用的Timeloop加速器建模框架,我们证明了新的优化方法可以显着改善先前的加速器设计的能量最小化和性能最大化。
The design space of possible schedules for mapping a Convolutional Neural Network layer onto a spatial accelerator array, referred as the dataflow, is enormous. The co-design of key architectural parameters (such as number of processing elements, sizes of register files and scratchpad memories) along with the dataflow to optimize the implementation of one or more CNN stages makes the design space explosively larger. Several recent efforts have addressed the design-space exploration problem for CNN accelerators via heuristics or limited search strategies. In this paper we develop the first optimization approach that uses analytical modeling and the solution of constrained nonlinear optimization problems for comprehensive algorithm-architecture co-design optimization. Using the Timeloop accelerator modeling framework, we demonstrate that the new optimization methodology can enable significant improvements over prior accelerator designs for both energy minimization and performance maximization.