Real-time neural network inference on extremely weak devices: agile offloading with explainable AI

Real-time neural network inference on extremely weak devices: agile offloading with explainable AI
复制标题

DOI:
10.1145/3495243.3560551
复制
发表时间:
2022-10
期刊:
Proceedings of the 28th Annual International Conference on Mobile Computing And Networking
影响因子:
--
通讯作者:
Kai Huang;Wei Gao
Kai Huang;Wei Gao
中科院分区:
其他
文献类型:
--
作者:
Kai Huang;Wei Gao

文献摘要

相似文献

随着人工智能应用的广泛应用,迫切需要在小型嵌入式设备上实现实时神经网络推理,但由于其能力极弱,在这些小型设备上部署神经网络并实现高性能的神经网络推理是具有挑战性的。尽管NN分区和卸载有助于此类部署,但它们无法最大限度地降低嵌入式设备的本地成本。相反,我们建议通过灵活的NN卸载来解决这一挑战,这将NN中所需的计算从在线推理迁移到离线学习。本文提出了一种新的神经网络卸载技术AgileNN,它利用可解释的人工智能技术在弱嵌入式设备上实现实时神经网络推理,从而在训练阶段显式地实施特征稀疏性,并最大限度地减少在线计算和通信成本。实验结果表明,AgileNN的推理延迟比现有方案降低了6倍,保证了嵌入式设备上的感知数据能够被及时消费。它还将本地设备的资源消耗减少到原来的1/8,而不会影响推理的准确性。
With the wide adoption of AI applications, there is a pressing need of enabling real-time neural network (NN) inference on small embedded devices, but deploying NNs and achieving high performance of NN inference on these small devices is challenging due to their extremely weak capabilities. Although NN partitioning and offloading can contribute to such deployment, they are incapable of minimizing the local costs at embedded devices. Instead, we suggest to address this challenge via agile NN offloading, which migrates the required computations in NN offloading from online inference to offline learning. In this paper, we present AgileNN, a new NN offloading technique that achieves real-time NN inference on weak embedded devices by leveraging eXplainable AI techniques, so as to explicitly enforce feature sparsity during the training phase and minimize the online computation and communication costs. Experiment results show that AgileNN's inference latency is >6X lower than the existing schemes, ensuring that sensory data on embedded devices can be timely consumed. It also reduces the local device's resource consumption by >8X, without impairing the inference accuracy.