Robust and Resource-efficient Machine Learning Aided Viewport Prediction in Virtual Reality

Robust and Resource-efficient Machine Learning Aided Viewport Prediction in Virtual Reality
复制标题

DOI:
10.1109/bigdata55660.2022.10020395
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Yuang Jiang;Konstantinos Poularakis;Diego Kiedanski;S. Kompella;L. Tassiulas
Yuang Jiang;Konstantinos Poularakis;Diego Kiedanski;S. Kompella;L. Tassiulas
中科院分区:
其他
文献类型:
--
作者:
Yuang Jiang;Konstantinos Poularakis;Diego Kiedanski;S. Kompella;L. Tassiulas

文献摘要

相似文献

360-近年来,由于头戴式显示器(HMD)和全景相机的快速发展,30度全景视频获得了相当大的关注。流式传输全景视频的一个主要问题是,与传统视频相比,全景视频的大小要大得多。此外,用户设备通常处于无线环境中,具有有限的电池、计算能力和带宽。为了减少资源消耗,研究人员提出了预测用户视口的方法,这样整个视频中只有一部分需要从服务器传输。然而,这种预测方法的鲁棒性在文献中被忽视了:通常假设只有少数几个模型,根据过去用户的经验进行预训练,适用于所有用户的预测。我们观察到,这些预先训练的模型对某些用户来说可能表现不佳,因为他们可能与大多数用户有着截然不同的行为,并且预先训练的模型无法捕捉到未见过的视频中的特征。在这项工作中,我们提出了一种新的基于Meta学习的视口预测范例,以减轻最差的预测性能,并确保视口预测的鲁棒性。该范例使用两个机器学习模型,其中第一个模型预测观看方向,第二个模型预测可以包括实际视口的最小视频预取大小。我们首先训练两个Meta模型,使它们对新的训练数据敏感,然后在用户观看视频时快速调整它们。评估结果表明,Meta模型可以快速适应每个用户,并可以显着提高预测精度,特别是对于性能最差的预测。
360-degree panoramic videos have gained considerable attention in recent years due to the rapid development of head-mounted displays (HMDs) and panoramic cameras. One major problem in streaming panoramic videos is that panoramic videos are much larger in size compared to traditional ones. Moreover, the user devices are often in a wireless environment, with limited battery, computation power, and bandwidth. To reduce resource consumption, researchers have proposed ways to predict the users’ viewports so that only part of the entire video needs to be transmitted from the server. However, the robustness of such prediction approaches has been overlooked in the literature: it is usually assumed that only a few models, pre-trained on past users’ experiences, are applied for prediction to all users. We observe that those pre-trained models can perform poorly for some users because they might have drastically different behaviors from the majority, and the pre-trained models cannot capture the features in unseen videos. In this work, we propose a novel meta learning based viewport prediction paradigm to alleviate the worst prediction performance and ensure the robustness of viewport prediction. This paradigm uses two machine learning models, where the first model predicts the viewing direction, and the second model predicts the minimum video prefetch size that can include the actual viewport. We first train two meta models so that they are sensitive to new training data, and then quickly adapt them to users while they are watching the videos. Evaluation results reveal that the meta models can adapt quickly to each user, and can significantly increase the prediction accuracy, especially for the worst-performing predictions.