Low-latency speculative inference on distributed multi-modal data streams

Low-latency speculative inference on distributed multi-modal data streams
复制标题

DOI:
10.1145/3458864.3467884
复制
发表时间:
2021-06
期刊:
Proceedings of the 19th Annual International Conference on Mobile Systems, Applications, and Services
影响因子:
--
通讯作者:
Tianxing Li;Jin Huang;Erik Risinger;Deepak Ganesan
Tianxing Li;Jin Huang;Erik Risinger;Deepak Ganesan
中科院分区:
其他
文献类型:
--
作者:
Tianxing Li;Jin Huang;Erik Risinger;Deepak Ganesan

文献摘要

相似文献

虽然多模态深度学习在人体跟踪、活动识别以及音频和视频分析等分布式传感任务中很有用,但在无线网络传感器系统中部署最先进的多模态模型带来了独特的挑战。不同模态的数据大小可能高度不对称(例如视频与音频),在存在无线动态的情况下,这些差异可能导致流之间出现显著的延迟。因此,一个缓慢的流可能会显著减慢云端的多模态推理系统,导致延迟增加(当被缓慢的流阻塞时)或者推理准确性下降(如果不等候就进行推理)。在本文中,我们引入对多模态数据流的推测性推理,以适应跨模态的这些不对称性。我们不是等到所有传感器流都到达并在时间上对齐才进行推理,而是对任何缺失、损坏或部分可用的传感器数据进行估算,然后使用学习到的模型和估算的数据进行推测性推理。回滚模块查看推测性推理的类别输出,并确定该类别对于不完整数据是否足够稳健以接受结果;如果不是,我们回滚推理并更新模型的输出。我们使用公开数据集在三个多模态应用场景中实现了该系统。实验结果表明,我们的系统在与六种最先进的方法具有相同准确性的情况下,实现了7 - 128倍的延迟加速。
While multi-modal deep learning is useful in distributed sensing tasks like human tracking, activity recognition, and audio and video analysis, deploying state-of-the-art multi-modal models in a wirelessly networked sensor system poses unique challenges. The data sizes for different modalities can be highly asymmetric (e.g., video vs. audio), and these differences can lead to significant delays between streams in the presence of wireless dynamics. Therefore, a slow stream can significantly slow down a multi-modal inference system in the cloud, leading to either increased latency (when blocked by the slow stream) or degradation in inference accuracy (if inference proceeds without waiting). In this paper, we introduce speculative inference on multi-modal data streams to adapt to these asymmetries across modalities. Rather than blocking inference until all sensor streams have arrived and been temporally aligned, we impute any missing, corrupt, or partially-available sensor data, then generate a speculative inference using the learned models and imputed data. A rollback module looks at the class output of speculative inference and determines whether the class is sufficiently robust to incomplete data to accept the result; if not, we roll back the inference and update the model's output. We implement the system in three multi-modal application scenarios using public datasets. The experimental results show that our system achieves 7 -- 128× latency speedup with the same accuracy as six state-of-the-art methods.