Self-Supervised Generation of Spatial Audio for 360 Video

Self-Supervised Generation of Spatial Audio for 360 Video
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Pedro Morgado;N. Vasconcelos;Timothy R. Langlois;Oliver Wang
Pedro Morgado;N. Vasconcelos;Timothy R. Langlois;Oliver Wang
中科院分区:
其他
文献类型:
--
作者:
Pedro Morgado;N. Vasconcelos;Timothy R. Langlois;Oliver Wang

文献摘要

相似文献

我们介绍了一种将360°摄像机录制的单声道音频转换为空间音频的方法,这是声音在整个观看球体上分布的表示。空间音频是沉浸式360°视频观看的重要组成部分,但在当前的360°视频制作中,空间音频麦克风仍然很少见。我们的系统由端到端可训练的神经网络组成,该网络可以分离单个声源并将其定位在观看范围内,条件是音频和360°视频帧的多模态分析。我们介绍了几个数据集,其中一个是我们自己拍摄的,另一个是从YouTube上收集的野外数据集,由360°视频上传和空间音频组成。在训练过程中,地面真实空间音频作为自我监督,混合单声道形成我们网络的输入。使用我们的方法,我们证明了仅基于同步360°视频和单声道音轨推断声音的空间定位是可能的。
We introduce an approach to convert mono audio recorded by a 360° video camera into spatial audio, a representation of the distribution of sound over the full viewing sphere. Spatial audio is an important component of immersive 360° video viewing, but spatial audio microphones are still rare in current 360° video production. Our system consists of end-to-end trainable neural networks that separate individual sound sources and localize them on the viewing sphere, conditioned on multi-modal analysis from the audio and 360° video frames. We introduce several datasets, including one filmed ourselves, and one collected in-the-wild from YouTube, consisting of 360° videos uploaded with spatial audio. During training, ground truth spatial audio serves as self-supervision and a mixed down mono track forms the input to our network. Using our approach we show that it is possible to infer the spatial localization of sounds based only on a synchronized 360° video and the mono audio track.