Are Existing Knowledge Transfer Techniques Effective for Deep Learning with Edge Devices?

Are Existing Knowledge Transfer Techniques Effective for Deep Learning with Edge Devices?
复制标题

DOI:
10.1109/edge.2018.00013
复制
发表时间:
2018-07
期刊:
2018 IEEE International Conference on Edge Computing (EDGE)
影响因子:
--
通讯作者:
Ragini Sharma;Saman Biookaghazadeh;Baoxin Li;Ming Zhao
Ragini Sharma;Saman Biookaghazadeh;Baoxin Li;Ming Zhao
中科院分区:
其他
文献类型:
--
作者:
Ragini Sharma;Saman Biookaghazadeh;Baoxin Li;Ming Zhao

文献摘要

被引文献

相似文献

随着边缘计算范式的出现,图像识别和增强现实等许多应用都需要在边缘设备上执行机器学习(ML)和人工智能(AI)任务。大多数人工智能和机器学习模型都很大,计算量很大,而边缘设备通常配备有限的计算和存储资源。为了在边缘设备上部署,这些模型可以被压缩和精简,但它们可能会失去功能,不能很好地执行。最近的研究使用知识转移技术将信息从一个大网络(称为教师)转移到一个小网络(称为学生),以提高后者的表现。这种方法似乎很有希望在边缘设备上学习,但缺乏对其有效性的彻底调查。本文对知识转移的性能(准确性和收敛速度)进行了广泛的研究,考虑了不同的学生体系结构和不同的教师向学生转移知识的技术。结果表明,KT的性能确实因体系结构和传输技术而异。通过将知识从教师的中间层和最后一层传递给较浅的学生,可以获得较好的绩效提升。但是其他架构和传输技术就没有那么好了,其中一些甚至会导致负面的性能影响。
With the emergence of edge computing paradigm, many applications such as image recognition and augmented reality require to perform machine learning (ML) and artificial intelligence (AI) tasks on edge devices. Most AI and ML models are large and computational-heavy, whereas edge devices are usually equipped with limited computational and storage resources. Such models can be compressed and reduced for deployment on edge devices, but they may lose their capability and not perform well. Recent works used knowledge transfer techniques to transfer information from a large network (termed teacher) to a small one (termed student) in order to improve the performance of the latter. This approach seems to be promising for learning on edge devices, but a thorough investigation on its effectiveness is lacking. This paper provides an extensive study on the performance (in both accuracy and convergence speed) of knowledge transfer, considering different student architectures and different techniques for transferring knowledge from teacher to student. The results show that the performance of KT does vary by architectures and transfer techniques. A good performance improvement is obtained by transferring knowledge from both the intermediate layers and last layer of the teacher to a shallower student. But other architectures and transfer techniques do not fare so well and some of them even lead to negative performance impact.