Acoustic Room Modelling Using 360 Stereo Cameras

Acoustic Room Modelling Using 360 Stereo Cameras
复制标题

DOI:
10.1109/tmm.2020.3037537
复制
发表时间:
2021
影响因子:
7.3
通讯作者:
Hansung Kim;Luca Remaggi;Sam Fowler;P. Jackson;A. Hilton
Hansung Kim;Luca Remaggi;Sam Fowler;P. Jackson;A. Hilton
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hansung Kim;Luca Remaggi;Sam Fowler;P. Jackson;A. Hilton

文献摘要

相似文献

在本文中,我们提出了一种利用球面360$^{\circ}$摄像机进行几何和属性预测的声学三维房间结构估计管道。代替设置麦克风阵列和扬声器来测量特定房间的声学参数,使用一对立体360摄像机简单实用的单镜头捕捉场景可以用来模拟这些声学参数。我们假设房间和物体可以表示为与房间坐标的主轴对齐的长方体(曼哈顿世界)。这个场景是用现成的消费者球形360相机作为立体对拍摄的。利用卷积神经网络(SegNet)对捕获的图像和语义标记进行对应匹配,估计出基于长方体的三维房间几何模型。估计的几何形状用于产生与场景频率相关的声学预测。据我们所知,这是文献中第一次尝试使用视觉几何估计和物体分类算法来预测声学特性。通过计算混响空间音频对象参数,将结果与测量结果进行比较,这些参数用于定制给定扬声器设置的混响再现。
In this paper we propose a pipeline for estimating acoustic 3D room structure with geometry and attribute prediction using spherical 360$^{\circ }$ cameras. Instead of setting microphone arrays with loudspeakers to measure acoustic parameters for specific rooms, a simple and practical single-shot capture of the scene using a stereo pair of 360 cameras can be used to simulate those acoustic parameters. We assume that the room and objects can be represented as cuboids aligned to the main axes of the room coordinate (Manhattan world). The scene is captured as a stereo pair using off-the-shelf consumer spherical 360 cameras. A cuboid-based 3D room geometry model is estimated by correspondence matching between captured images and semantic labelling using a convolutional neural network (SegNet). The estimated geometry is used to produce frequency-dependent acoustic predictions of the scene. This is, to our knowledge, the first attempt in the literature to use visual geometry estimation and object classification algorithms to predict acoustic properties. Results are compared to measurements through calculated reverberant spatial audio object parameters used for reverberation reproduction customized to the given loudspeaker set up.