Few-Shot Audio-Visual Learning of Environment Acoustics

Few-Shot Audio-Visual Learning of Environment Acoustics
复制标题

DOI:
10.48550/arxiv.2206.04006
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Sagnik Majumder;Changan Chen;Ziad Al-Halah;K. Grauman
Sagnik Majumder;Changan Chen;Ziad Al-Halah;K. Grauman
中科院分区:
其他
文献类型:
--
作者:
Sagnik Majumder;Changan Chen;Ziad Al-Halah;K. Grauman

文献摘要

被引文献

相似文献

房间脉冲响应(RIR)功能捕捉周围物理环境如何改变听者听到的声音,这对AR、VR和机器人技术的各种应用都有影响。传统估计rir的方法假设在整个环境中进行密集的几何和/或声音测量,而我们探索如何基于空间中观察到的稀疏图像和回声集来推断rir。为了实现这一目标,我们引入了一种基于变压器的方法,该方法利用自关注来构建丰富的声学环境,然后通过交叉关注来预测任意查询源接收器位置的rir。此外,我们设计了一个新的训练目标,提高了RIR预测与目标之间的声学特征匹配。在使用最先进的3D环境视听模拟器的实验中,我们证明了我们的方法成功地生成了任意的rir,优于最先进的方法,并且-与传统方法有很大的不同-以少数镜头的方式推广到新环境。项目:http://vision.cs.utexas.edu/projects/fs_rir。
Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional methods to estimate RIRs assume dense geometry and/or sound measurements throughout the environment, we explore how to infer RIRs based on a sparse set of images and echoes observed in the space. Towards that goal, we introduce a transformer-based method that uses self-attention to build a rich acoustic context, then predicts RIRs of arbitrary query source-receiver locations through cross-attention. Additionally, we design a novel training objective that improves the match in the acoustic signature between the RIR predictions and the targets. In experiments using a state-of-the-art audio-visual simulator for 3D environments, we demonstrate that our method successfully generates arbitrary RIRs, outperforming state-of-the-art methods and -- in a major departure from traditional methods -- generalizing to novel environments in a few-shot manner. Project: http://vision.cs.utexas.edu/projects/fs_rir.