Exploring Spherical Autoencoder for Spherical Video Content Processing

Exploring Spherical Autoencoder for Spherical Video Content Processing
复制标题

DOI:
10.1145/3503161.3548364
复制
发表时间:
2022-10
期刊:
Proceedings of the 30th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
Jin Zhou;Na Li;Yao Liu-;Shuochao Yao;Songqing Chen
Jin Zhou;Na Li;Yao Liu-;Shuochao Yao;Songqing Chen
中科院分区:
其他
文献类型:
--
作者:
Jin Zhou;Na Li;Yao Liu-;Shuochao Yao;Songqing Chen

文献摘要

相似文献

3D球形内容越来越多地出现在各种应用(例如增强现实/混合现实/虚拟现实)中,以提供更好的用户沉浸感体验,然而如今对这种球形3D内容的处理在投影后仍然主要依赖传统的2D方法,导致关键信息的失真和/或丢失。本研究旨在探索更直接且更有效地处理球形3D内容的方法。以360度视频为例,我们提出一种名为球形自动编码器(SAE)的新方法用于球形视频处理。SAE不是投影到2D空间,而是将360度视频内容表示为一个球形物体,并直接对360度视频进行编码和解码。此外,为了支持在通常资源受限的普及型移动设备上采用SAE,我们在SAE的基础上进一步提出了两种优化方法。首先,由于视场角(FoV)预测已被广泛研究并用于仅将部分内容传输到移动设备以节省带宽和电池消耗,我们设计了p - SAE,这是一种具有部分视图支持的SAE方案,可以利用这种FoV预测。其次,由于机器学习模型在移动设备上运行时经常被压缩以减少处理负载,这通常会导致输出质量下降(例如SAE中的视频质量),我们通过将压缩感知理论应用于SAE提出了c - SAE,以便在模型被压缩时保持视频质量。我们大量的实验表明,直接合并和处理球形信号是有前景的,并且它在很大程度上优于传统方法。p - SAE和c - SAE在单独使用或与模型压缩一起使用时,在提供高质量视频(例如峰值信噪比结果)方面都显示出了它们的有效性。
3D spherical content is increasingly presented in various applications (e.g., AR/MR/VR) for better users' immersiveness experience, yet today processing such spherical 3D content still mainly relies on the traditional 2D approaches after projection, leading to the distortion and/or loss of critical information. This study sets to explore methods to process spherical 3D content directly and more effectively. Using 360-degree videos as an example, we propose a novel approach called Spherical Autoencoder (SAE) for spherical video processing. Instead of projecting to a 2D space, SAE represents the 360-degree video content as a spherical object and employs encoding and decoding on the 360-degree video directly. Furthermore, to support the adoption of SAE on pervasive mobile devices that often have resource constraints, we further propose two optimizations on top of SAE.First, since the FoV (Field of View) prediction is widely studied and leveraged to transport only a portion of the content to the mobile device to save bandwidth and battery consumption, we design p-SAE, a SAE scheme with the partial view support that can utilize such FoV prediction. Second, since machine learning models are often compressed when running on mobile devices in order to reduce the processing load, which usually leads to degradation of output (e.g., video quality in SAE), we propose c-SAE by applying the compressive sensing theory into SAE to maintain the video quality when the model is compressed. Our extensive experiments show that directly incorporating and processing spherical signals is promising, and it outperforms the traditional approaches by a large margin. Both p-SAE and c-SAE show their effectiveness in delivering high quality videos (e.g., PSNR results) when used alone or combined together with model compression.