14.7 A 288µW programmable deep-learning processor with 270KB on-chip weight storage using non-uniform memory hierarchy for mobile intelligence

14.7 A 288µW programmable deep-learning processor with 270KB on-chip weight storage using non-uniform memory hierarchy for mobile intelligence
复制标题

DOI:
10.1109/isscc.2017.7870355
复制
发表时间:
2017-02
期刊:
2017 IEEE International Solid-State Circuits Conference (ISSCC)
影响因子:
--
通讯作者:
Suyoung Bang;Jingcheng Wang;Ziyun Li;Cao Gao;Yejoong Kim;Qing Dong;Yen-Po Chen;Laura Fick;Xun Sun;R. Dreslinski;T. Mudge;Hun-Seok Kim;D. Blaauw;D. Sylvester
Suyoung Bang;Jingcheng Wang;Ziyun Li;Cao Gao;Yejoong Kim;Qing Dong;Yen-Po Chen;Laura Fick;Xun Sun;R. Dreslinski;T. Mudge;Hun-Seok Kim;D. Blaauw;D. Sylvester
中科院分区:
其他
文献类型:
--
作者:
Suyoung Bang;Jingcheng Wang;Ziyun Li;Cao Gao;Yejoong Kim;Qing Dong;Yen-Po Chen;Laura Fick;Xun Sun;R. Dreslinski;T. Mudge;Hun-Seok Kim;D. Blaauw;D. Sylvester

文献摘要

被引文献

相似文献

深度学习已被证明是广泛应用的强大工具,例如语音识别和对象检测等。最近,人们对移动的物联网[1]的深度学习越来越感兴趣,以便在边缘实现智能,并通过仅转发有意义的事件来保护云免受大量数据的影响。这种分层智能因此通过在边缘设备处权衡计算和通信来增强无线电带宽和功率效率。由于许多移动的应用是“永远在线”的(例如,语音命令),低功率是关键的设计约束。然而,之前的工作集中在高性能可重构处理器[2-3]上,针对消耗> 50 mW的大规模深度神经网络(DNN)进行了优化。DRAM中的片外权重存储在现有技术中也很常见[2-3],这意味着由于密集的片外数据移动而导致的显著的额外功耗。
Deep learning has proven to be a powerful tool for a wide range of applications, such as speech recognition and object detection, among others. Recently there has been increased interest in deep learning for mobile IoT [1] to enable intelligence at the edge and shield the cloud from a deluge of data by only forwarding meaningful events. This hierarchical intelligence thereby enhances radio bandwidth and power efficiency by trading-off computation and communication at edge devices. Since many mobile applications are “always-on” (e.g., voice commands), low power is a critical design constraint. However, prior works have focused on high performance reconfigurable processors [2–3] optimized for large-scale deep neural networks (DNNs) that consume >50mW. Off-chip weight storage in DRAM is also common in the prior works [2–3], which implies significant additional power consumption due to intensive off-chip data movement.