How to Obtain and Run Light and Efficient Deep Learning Networks

How to Obtain and Run Light and Efficient Deep Learning Networks
复制标题

DOI:
10.1109/iccad45719.2019.8942106
复制
发表时间:
2019-11
期刊:
2019 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)
影响因子:
--
通讯作者:
Fan Chen;W. Wen;Linghao Song;Jingchi Zhang;H. Li;Yiran Chen
Fan Chen;W. Wen;Linghao Song;Jingchi Zhang;H. Li;Yiran Chen
中科院分区:
其他
文献类型:
--
作者:
Fan Chen;W. Wen;Linghao Song;Jingchi Zhang;H. Li;Yiran Chen

文献摘要

相似文献

随着深度神经网络 (DNN) 的模型规模不断增大以获得更好的性能,与训练和测试相关的计算成本的增加使得在资源有限的终端/边缘设备上部署 DNN 变得极其困难,同时又满足响应时间要求。为了应对这一挑战,深度学习社会广泛采用模型压缩来压缩模型大小,从而降低计算成本。然而,在这些算法级解决方案中,硬件设计的实际影响往往被忽视,例如对内存层次结构的随机访问的增加和内存容量的限制。另一方面,对算法级别的计算需求的有限理解可能会导致在硬件设计过程中做出不切实际的假设。在这项工作中,我们将讨论这种不匹配,并提供我们的方法如何通过跨软件和硬件级别的交互式设计实践来解决它。
As the model size of deep neural networks (DNNs) grows for better performance, the increase in computational cost associated with training and testing makes it extremely difficulty to deploy DNNs on end/edge devices with limited resources while also satisfying the response time requirement. To address this challenge, model compression which compresses model size and thus reduces computation cost is widely adopted in deep learning society. However, the practical impacts of hardware design are often ignored in these algorithm-level solutions, such as the increase of the random accesses to memory hierarchy and the constraints of memory capacity. On the other side, limited understanding about the computational needs at algorithm level may lead to unrealistic assumptions during the hardware designs. In this work, we will discuss this mismatch and provide how our approach addresses it through an interactive design practice across both software and hardware levels.