LiveObj: Object Semantics-based Viewport Prediction for Live Mobile Virtual Reality Streaming

LiveObj: Object Semantics-based Viewport Prediction for Live Mobile Virtual Reality Streaming
复制标题

DOI:
10.1109/tvcg.2021.3067686
复制
发表时间:
2021-04
影响因子:
5.2
通讯作者:
Xianglong Feng;Zeyang Bao;Sheng Wei
Xianglong Feng;Zeyang Bao;Sheng Wei
中科院分区:
计算机科学1区
文献类型:
--
作者:
Xianglong Feng;Zeyang Bao;Sheng Wei

文献摘要

相似文献

虚拟现实(VR)视频流(又称360度视频流)作为一种新的多媒体形式,为用户提供身临其境的观看体验,近年来越来越受到人们的欢迎。然而,360度视频帧的大量数据带来了巨大的带宽挑战。已经做出了研究努力,通过预测和选择性地流传输用户的视窗来降低带宽消耗。然而,现有的方法需要历史用户或视频数据,不能应用于直播这一最具吸引力的VR流媒体场景。本文提出了一种实时视点预测机制LiveObj,该机制通过对视频中的对象进行语义检测来实现。然后,通过使用强化学习算法来跟踪检测到的对象以实时推断用户的视区。我们基于48个用户观看了10个VR视频的评估结果表明,LiveObj具有很高的预测精度和显著的带宽节约。此外,LiveObj在低处理延迟下实现了实时性能,满足了VR直播的要求。
Virtual reality (VR) video streaming (a.k.a., 360-degree video streaming) has been gaining popularity recently as a new form of multimedia providing the users with immersive viewing experience. However, the high volume of data for the 360-degree video frames creates significant bandwidth challenges. Research efforts have been made to reduce the bandwidth consumption by predicting and selectively streaming the user's viewports. However, the existing approaches require historical user or video data and cannot be applied to live streaming, the most attractive VR streaming scenario. We develop a live viewport prediction mechanism, namely LiveObj, by detecting the objects in the video based on their semantics. The detected objects are then tracked to infer the user's viewport in real time by employing a reinforcement learning algorithm. Our evaluations based on 48 users watching 10 VR videos demonstrate high prediction accuracy and significant bandwidth savings obtained by LiveObj. Also, LiveObj achieves real-time performance with low processing delays, meeting the requirement of live VR streaming.