Move2Hear: Active Audio-Visual Source Separation

Move2Hear: Active Audio-Visual Source Separation
复制标题

DOI:
10.1109/iccv48922.2021.00034
复制
发表时间:
2021-05
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Sagnik Majumder;Ziad Al-Halah;K. Grauman
Sagnik Majumder;Ziad Al-Halah;K. Grauman
中科院分区:
其他
文献类型:
--
作者:
Sagnik Majumder;Ziad Al-Halah;K. Grauman

文献摘要

被引文献

相似文献

我们介绍了活跃的视听源分离问题,代理必须聪明地移动,以更好地隔离来自其环境中感兴趣的对象的声音。引入一种强化学习方法,该方法在改进预测的音频分离质量的指导下,训练了控制代理人的相机和麦克风放置的行动。展示了我们模型的最小运动序列的能力,该序列具有音频源分离的最大收益。
We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple audio sources simultaneously (e.g., a person speaking down the hall in a noisy household) and it must use its eyes and ears to automatically separate out the sounds originating from a target object within a limited time budget. Towards this goal, we introduce a reinforcement learning approach that trains movement policies controlling the agent’s camera and microphone placement over time, guided by the improvement in predicted audio separation quality. We demonstrate our approach in scenarios motivated by both augmented reality (system is already co-located with the target object) and mobile robotics (agent begins arbitrarily far from the target object). Using state-of-the-art realistic audio-visual simulations in 3D environments, we demonstrate our model’s ability to find minimal movement sequences with maximal payoff for audio source separation. Project: http://vision.cs.utexas.edu/projects/move2hear.