Edge-MultiAI: Multi-Tenancy of Latency-Sensitive Deep Learning Applications on Edge

Edge-MultiAI: Multi-Tenancy of Latency-Sensitive Deep Learning Applications on Edge
复制标题

DOI:
10.1109/ucc56403.2022.00012
复制
发表时间:
2022-11
期刊:
2022 IEEE/ACM 15th International Conference on Utility and Cloud Computing (UCC)
影响因子:
--
通讯作者:
S. Zobaed;Ali Mokhtari;J. Champati;M. Kourouma;M. Salehi
S. Zobaed;Ali Mokhtari;J. Champati;M. Kourouma;M. Salehi
中科院分区:
其他
文献类型:
--
作者:
S. Zobaed;Ali Mokhtari;J. Champati;M. Kourouma;M. Salehi

文献摘要

被引文献

相似文献

智能物联网系统通常需要连续执行多个延迟敏感的深度学习(DL)应用程序。边缘服务器充当这种基于IoT的系统的基石,然而,它们的资源限制阻碍了多个(多租户)DL应用的连续执行。挑战在于,深度学习应用程序基于庞大的“神经网络(NN)模型”运行,这些模型无法同时在边缘的有限内存空间中维护。因此,本研究的主要贡献是克服内存争用的挑战,从而满足延迟限制的DL应用程序,而不影响他们的推理精度。我们提出了一个高效的NN模型管理框架,称为Edge-MultiAI,它将DL应用程序的NN模型引入边缘内存,从而最大化多租户程度和热启动次数。Edge-MultiAI利用NN模型压缩技术,例如模型量化,并动态加载DL应用程序的NN模型,以刺激边缘服务器上的多租户。我们还为Edge-MultiAI设计了一个模型管理启发式算法,称为iWS-BFE,它基于贝叶斯理论来预测多租户应用程序的推理请求,并使用它来选择合适的NN模型进行加载,从而增加热启动推理的数量。我们评估了Edge-MultiAI在各种配置下的有效性和鲁棒性。结果表明,Edge-MultiAI可以将边缘上的多租户程度刺激至少2倍,并将热启动次数增加60%,而不会对应用程序的推理准确性造成任何重大损失。
Smart IoT-based systems often desire continuous execution of multiple latency-sensitive Deep Learning (DL) applications. The edge servers serve as the cornerstone of such IoT based systems, however, their resource limitations hamper the continuous execution of multiple (multi-tenant) DL applications. The challenge is that, DL applications function based on bulky “neural network (NN) models” that cannot be simultaneously maintained in the limited memory space of the edge. Accordingly, the main contribution of this research is to overcome the memory contention challenge, thereby, meeting the latency constraints of the DL applications without compromising their inference accuracy. We propose an efficient NN model management framework, called Edge-MultiAI, that ushers the NN models of the DL applications into the edge memory such that the degree of multi-tenancy and the number of warm-starts are maximized. Edge-MultiAI leverages NN model compression techniques, such as model quantization, and dynamically loads NN models for DL applications to stimulate multi-tenancy on the edge server. We also devise a model management heuristic for Edge-MultiAI, called iWS-BFE, that functions based on the Bayesian theory to predict the inference requests for multi-tenant applications, and uses it to choose the appropriate NN models for loading, hence, increasing the number of warm-start inferences. We evaluate the efficacy and robustness of Edge-MultiAI under various configurations. The results reveal that Edge-MultiAI can stimulate the degree of multi-tenancy on the edge by at least 2× and increase the number of warm-starts by ≈60% without any major loss on the inference accuracy of the applications.