STI: Turbocharge NLP Inference at the Edge via Elastic Pipelining

STI: Turbocharge NLP Inference at the Edge via Elastic Pipelining
复制标题

DOI:
10.1145/3575693.3575698
复制
发表时间:
2022-07
期刊:
Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2
影响因子:
--
通讯作者:
Liwei Guo;Wonkyo Choe;F. Lin
Liwei Guo;Wonkyo Choe;F. Lin
中科院分区:
其他
文献类型:
--
作者:
Liwei Guo;Wonkyo Choe;F. Lin

文献摘要

相似文献

移动应用越来越多地采用自然语言处理(NLP)推理,在移动应用中,设备上的推理对于保护用户数据隐私和避免网络往返至关重要。然而,NLP模型的空前大小同时强调延迟和内存,从而在移动设备的两个关键资源之间产生紧张关系。为了满足目标延迟,将整个模型保存在内存中可以尽快启动执行,但会增加一个应用程序的内存占用数倍,在被移动内存管理回收之前,将其优势限制在几个推断中。另一方面,按需从存储中加载模型会导致长达几秒钟的IO,远远超过用户满意的延迟范围;由于IO和计算延迟之间的高度偏度,流水线分层模型加载和执行也不会隐藏IO。为此,我们提出了快速变压器推理(STI)。STI建立在最大化模型最重要部分的IO/计算资源利用率的关键思想之上,通过两种新技术来协调延迟与内存紧张。首先,模型分片。STI将模型参数作为独立可调的分片进行管理,并分析它们对准确性的重要性。二是弹性管道规划带预载缓冲器。STI实例化一个IO/计算管道,并为预加载碎片使用一个小缓冲区来引导执行,而不会在早期阶段停滞;它根据资源弹性执行的重要性明智地选择、调整和组装碎片,最大限度地提高推理准确性。在两个商用soc之上,我们构建STI并根据广泛的NLP任务,在实际的目标延迟范围内,以及CPU和GPU上对其进行评估。我们证明STI在低内存1-2个数量级的情况下提供了高精度,优于竞争基准。
Natural Language Processing (NLP) inference is seeing increasing adoption by mobile applications, where on-device inference is desirable for crucially preserving user data privacy and avoiding network roundtrips. Yet, the unprecedented size of an NLP model stresses both latency and memory, creating a tension between the two key resources of a mobile device. To meet a target latency, holding the whole model in memory launches execution as soon as possible but increases one app’s memory footprints by several times, limiting its benefits to only a few inferences before being recycled by mobile memory management. On the other hand, loading the model from storage on demand incurs IO as long as a few seconds, far exceeding the delay range satisfying to a user; pipelining layerwise model loading and execution does not hide IO either, due to the high skewness between IO and computation delays. To this end, we propose Speedy Transformer Inference (STI). Built on the key idea of maximizing IO/compute resource utilization on the most important parts of a model, STI reconciles the latency v.s. memory tension via two novel techniques. First, model sharding. STI manages model parameters as independently tunable shards, and profiles their importance to accuracy. Second, elastic pipeline planning with a preload buffer. STI instantiates an IO/compute pipeline and uses a small buffer for preload shards to bootstrap execution without stalling at early stages; it judiciously selects, tunes, and assembles shards per their importance for resource-elastic execution, maximizing inference accuracy. Atop two commodity SoCs, we build STI and evaluate it against a wide range of NLP tasks, under a practical range of target latencies, and on both CPU and GPU. We demonstrate that STI delivers high accuracies with 1–2 orders of magnitude lower memory, outperforming competitive baselines.