Secure and Efficient Mobile DNN Using Trusted Execution Environments

Secure and Efficient Mobile DNN Using Trusted Execution Environments
复制标题

DOI:
10.1145/3579856.3582820
复制
发表时间:
2023-07
期刊:
Proceedings of the 2023 ACM Asia Conference on Computer and Communications Security
影响因子:
--
通讯作者:
B. Hu;Yan Wang;Jerry Q. Cheng;Tianming Zhao;Yucheng Xie;Xiaonan Guo;Ying Chen
B. Hu;Yan Wang;Jerry Q. Cheng;Tianming Zhao;Yucheng Xie;Xiaonan Guo;Ying Chen
中科院分区:
其他
文献类型:
--
作者:
B. Hu;Yan Wang;Jerry Q. Cheng;Tianming Zhao;Yucheng Xie;Xiaonan Guo;Ying Chen

文献摘要

相似文献

许多移动的应用程序都采用了深度神经网络(DNN),因为它们具有强大的推理能力。由于输入数据和DNN架构都可能是敏感的,因此对在移动的设备上安全执行DNN的需求越来越大。为此,移动的设备(移动的TEE)上的基于硬件的可信执行环境(诸如ARM TrustZone)最近被利用来安全地执行CNN。然而,在移动的TEE上运行整个DNN是具有挑战性的,因为TEE具有严格的资源和性能约束。在这项工作中,我们开发了一种新的基于移动的TEE的安全框架,可以有效地执行整个DNN在资源受限的移动的TEE与最小的推理时间开销。具体来说,我们提出了一种渐进式修剪,以逐渐识别和删除DNN中的冗余神经元,同时保持较高的推理精度。接下来,我们开发了一种内存优化方法,利用低级编程技术重新分配修剪神经元的内存存储。最后,我们设计了一种新的自适应分区方法,该方法根据移动的TEE中的可用内存将修剪后的模型划分为多个分区,并以最小的加载时间开销将分区分别加载到移动的TEE中。我们对各种DNN和开源数据集的实验表明,与使用移动的TEE保护整个DNN的现有方法相比,我们可以实现2-30倍的推理时间,并且具有相当的准确性。
Many mobile applications have resorted to deep neural networks (DNNs) because of their strong inference capabilities. Since both input data and DNN architectures could be sensitive, there is an increasing demand for secure DNN execution on mobile devices. Towards this end, hardware-based trusted execution environments on mobile devices (mobile TEEs), such as ARM TrustZone, have recently been exploited to execute CNN securely. However, running entire DNNs on mobile TEEs is challenging as TEEs have stringent resource and performance constraints. In this work, we develop a novel mobile TEE-based security framework that can efficiently execute the entire DNN in a resource-constrained mobile TEE with minimal inference time overhead. Specifically, we propose a progressive pruning to gradually identify and remove the redundant neurons from a DNN while maintaining a high inference accuracy. Next, we develop a memory optimization method to deallocate the memory storage of the pruned neurons utilizing the low-level programming technique. Finally, we devise a novel adaptive partitioning method that divides the pruned model into multiple partitions according to the available memory in the mobile TEE and loads the partitions into the mobile TEE separately with a minimal loading time overhead. Our experiments with various DNNs and open-source datasets demonstrate that we can achieve 2-30 times less inference time with comparable accuracy compared to existing approaches securing entire DNNs with mobile TEE.