Understanding the potential of server-driven edge video analytics

Understanding the potential of server-driven edge video analytics
复制标题

DOI:
10.1145/3508396.3512872
复制
发表时间:
2022-03
期刊:
Proceedings of the 23rd Annual International Workshop on Mobile Computing Systems and Applications
影响因子:
--
通讯作者:
Qizheng Zhang;Kuntai Du;Neil Agarwal;R. Netravali;Junchen Jiang
Qizheng Zhang;Kuntai Du;Neil Agarwal;R. Netravali;Junchen Jiang
中科院分区:
其他
文献类型:
--
作者:
Qizheng Zhang;Kuntai Du;Neil Agarwal;R. Netravali;Junchen Jiang

文献摘要

相似文献

边缘视频分析应用的激增催生了一种新的流媒体协议,该协议将积极压缩的视频流传输到远程服务器,以进行计算密集型DNN推理。这种协议的一个流行设计范例是利用服务器端DNN来提取有用的反馈(例如,基于发送到服务器的低质量编码流),并使用反馈来通知摄像机将来应该如何编码和流式传输视频。在这种服务器驱动的方法中,反馈的理想形式应当(1)从来自视频传感器的最少信息导出(2)招致最小带宽使用以获得(3)指示最优视频流传输/编码方案(例如,需要高编码质量的最少帧/区域)。然而,我们的初步研究表明,这些理想化的要求远远没有得到满足。使用对象检测作为一个示例用例,我们通过考虑更广泛的设计空间,在如何从DNN获得反馈,应该多久提取一次,以及如何确定我们提取反馈的视频的编码质量方面,展示了显著但尚未开发的改进空间。
The proliferation of edge video analytics applications has given rise to a new breed of streaming protocols which stream aggressively compressed videos to remote servers for compute-intensive DNN inference. One popular design paradigm of such protocols is to leverage the server-side DNN to extract useful feedback (e.g. based on a low-quality-encoded stream sent to the server) and use the feedback to inform how the camera should encode and stream the video in the future. In this server-driven approach, an ideal form of feedback should (1) be derived from minimum information from the video sensor (2) incur minimum bandwidth usage to obtain (3) indicate the optimal video streaming/encoding scheme (e.g. the minimum frames/regions that require high encoding quality). However, our preliminary study shows that these idealized requirements are far from being met. Using object detection as an example use case, we demonstrate significant yet untapped room for improvement by considering a broader design space, in terms of how the feedback should be derived from the DNN, how often it should be extracted, and how to determine the encoding quality of the video on which we extract the feedback.