LiveDeep: Online Viewport Prediction for Live Virtual Reality Streaming Using Lifelong Deep Learning

LiveDeep: Online Viewport Prediction for Live Virtual Reality Streaming Using Lifelong Deep Learning
复制标题

LiveDeep:使用终身深度学习进行实时虚拟现实流的在线视口预测

DOI:
10.1109/vr46266.2020.00104
复制
发表时间:
2020
期刊:
IEEE Conference on Virtual Reality and 3D User Interfaces (VR
影响因子:
--
通讯作者:
Wei, Sheng
Wei, Sheng
中科院分区:
--
文献类型:
--
作者:
Feng, Xianglong;Liu, Yao;Wei, Sheng

文献摘要

参考文献

相似文献

直播虚拟现实 (VR) 流媒体已成为消费市场中流行且趋势的视频应用,为用户提供 360 度、身临其境的观看体验。为了提供优质的体验,由于带宽消耗显着增加,VR 流媒体面临着独特的挑战。为了解决带宽挑战,VR 视频视口预测被提出作为一种可行的解决方案,它仅预测用户感兴趣的视口并将其高质量传输到 VR 设备。然而,大多数现有的视口预测方法仅针对视频点播(VOD)用例,需要对实时流场景中不可用的历史视频和/或用户数据进行离线处理。在这项工作中,我们开发了一种用于实时 VR 流媒体的新颖视口预测方法,该方法只需要当前观看会话中的视频内容和用户数据。为了解决训练数据不足和实时处理的挑战,我们提出了一种实时VR专用的深度学习机制,即LiveDeep,来创建在线视口预测模型并进行实时推理。 LiveDeep 采用混合方法来解决实时 VR 流媒体中的独特挑战,包括 (1) 备用在线数据收集、标记、训练和推理计划,以及受控反馈循环以适应稀疏的训练数据; (2)混合混合神经网络模型,以适应单一模型引起的不准确性。我们使用从公共 VR 用户头部运动数据集中获得的 48 个用户和 14 个不同类型的 VR 视频来评估 LiveDeep。结果表明,预测准确率约为 90%,带宽节省约 40%,处理时间较长,满足 VR 直播的带宽和实时性要求。
Live virtual reality (VR) streaming has become a popular and trending video application in the consumer market providing users with 360-degree, immersive viewing experiences. To provide premium quality of experience, VR streaming faces unique challenges due to the significantly increased bandwidth consumption. To address the bandwidth challenge, VR video viewport prediction has been proposed as a viable solution, which predicts and streams only the user’s viewport of interest with high quality to the VR device. However, most of the existing viewport prediction approaches target only the video-on-demand (VOD) use cases, requiring offline processing of the historical video and/or user data that are not available in the live streaming scenario. In this work, we develop a novel viewport prediction approach for live VR streaming, which only requires video content and user data in the current viewing session. To address the challenges of insufficient training data and real-time processing, we propose a live VR-specific deep learning mechanism, namely LiveDeep, to create the online viewport prediction model and conduct real-time inference. LiveDeep employs a hybrid approach to address the unique challenges in live VR streaming, involving (1) an alternate online data collection, labeling, training, and inference schedule with controlled feedback loop to accommodate for the sparse training data; and (2) a mixture of hybrid neural network models to accommodate for the inaccuracy caused by a single model. We evaluate LiveDeep using 48 users and 14 VR videos of various types obtained from a public VR user head movement dataset. The results indicate around 90% prediction accuracy, around 40% bandwidth savings, and premium processing time, which meets the bandwidth and real-time requirements of live VR streaming.
DOI: 10.1145/2578260.2578277
发表时间: 2014-03
期刊: Proceedings of Network and Operating System Support on Digital Audio and Video Workshop
影响因子: --
作者:
Sheng Wei;Viswanathan Swaminathan
通讯作者: Sheng Wei;Viswanathan Swaminathan
使用 HTTP/2 改进虚拟现实流
DOI: 10.1145/3083187.3083224
发表时间: 2017
期刊: Proceedings of the 8th ACM on Multimedia Systems Conference
影响因子: --
作者:
Stefano Petrangeli;F. Turck;Viswanathan Swaminathan;Mohammad Hosseini
通讯作者: Mohammad Hosseini
DOI: 10.1145/3204949.3208119
发表时间: 2018
期刊: Proceedings of the 9th ACM Multimedia Systems Conference
影响因子: --
作者:
Jangwoo Son;Dongmin Jang;Eun‐Seok Ryu
通讯作者: Eun‐Seok Ryu
DOI: 10.5121/ijcseit.2013.3503
发表时间: 2013
影响因子: 23.6
作者:
G. Sivakumar;V. Venkatachalam;PG. Scholar;Erode Sengunthar Engg
通讯作者: Erode Sengunthar Engg