Reconstructing room scales with a single sound for augmented reality displays

Reconstructing room scales with a single sound for augmented reality displays
复制标题

使用单一声音重建房间尺度以进行增强现实显示

DOI:
10.1080/15980316.2022.2145377
复制
发表时间:
2023
影响因子:
3.7
通讯作者:
Sun, Qi
Sun, Qi
中科院分区:
工程技术3区
文献类型:
--
作者:
Liang, Benjamin S.;Liang, Andrew S.;Roman, Iran;Weiss, Tomer;Duinkharjav, Budmonde;Bello, Juan Pablo;Sun, Qi

文献摘要

参考文献

被引文献

相似文献

感知和重建我们的3D物理环境是增强现实(AR)显示器的一项重要任务,具有广泛的应用。例如,重建的几何形状通常用于在精确位置处显示3D对象。虽然相机捕获的图像是用于逼真地重建3D物理环境的常用数据源,但它们仅限于视线环境,需要耗时且重复的数据捕获技术来捕获完整的3D图片。例如,当前的AR设备需要用户扫描整个房间以获得其几何尺寸。当空间被遮挡或不可接近时,这种光学过程是乏味的且不适用的。音频波通过从不同表面反弹而在空间中传播,但不像光那样被单个物体(如墙壁)“遮挡”。在这项研究中,我们的目标是问一个问题:“一个人能听到房间的大小吗?”为了回答这个问题,我们提出了一种仅从单个声音推断房间几何形状的方法,我们将其定义为从单个扬声器播放的音频波序列,利用深度学习来解码来自单个扬声器和麦克风系统的隐含空间信息。通过一系列的实验和研究,我们的工作证明了我们的方法在推断三维环境的空间布局的有效性。我们的工作在多模态布局重构中引入了一个鲁棒的构建块。
Perception and reconstruction of our 3D physical environment is an essential task with broad applications for Augmented Reality (AR) displays. For example, reconstructed geometries are commonly leveraged for displaying 3D objects at accurate positions. While camera-captured images are a frequently used data source for realistically reconstructing 3D physical surroundings, they are limited to line-of-sight environments, requiring time-consuming and repetitive data-capture techniques to capture a full 3D picture. For instance, current AR devices require users to scan through a whole room to obtain its geometric sizes. This optical process is tedious and inapplicable when the space is occluded or inaccessible. Audio waves propagate through space by bouncing from different surfaces, but are not 'occluded' by a single object such as a wall, unlike light. In this research, we aim to ask the question‘can one hear the size of a room?’. To answer that, we propose an approach for inferring room geometries only from a single sound, which we define as an audio wave sequence played from a single loud speaker, leveraging deep learning for decoding implicitly-carried spatial information from a single speaker-and-microphone system. Through a series of experiments and studies, our work demonstrates our method's effectiveness at inferring a 3D environment's spatial layout. Our work introduces a robust building block in multi-modal layout reconstruction.
DOI: 10.1109/tmm.2015.2428998
发表时间: 2015-10-01
影响因子: 7.3
作者:
Stowell, Dan;Giannoulis, Dimitrios;Plumbley, Mark D.
通讯作者: Plumbley, Mark D.