Complementary Cues from Audio Help Combat Noise in Weakly-Supervised Object Detection

Complementary Cues from Audio Help Combat Noise in Weakly-Supervised Object Detection
复制标题

DOI:
10.1109/wacv56688.2023.00222
复制
发表时间:
2023-01
期刊:
2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Cagri Gungor;Adriana Kovashka
Cagri Gungor;Adriana Kovashka
中科院分区:
其他
文献类型:
--
作者:
Cagri Gungor;Adriana Kovashka

文献摘要

相似文献

我们解决了在噪声环境中学习对象检测器的问题,这是弱监督学习的重大挑战之一。我们使用多模式学习来帮助定位感兴趣的对象,但与其他方法不同的是,我们将音频视为一种辅助模式,帮助处理从视觉区域检测中的噪声。首先,我们使用视听模型为训练集生成新的“地面真实”标签,以消除视觉特征和噪声监督之间的噪声。其次,我们提出了一条音频和类别预测之间的间接路径,它结合了视觉和音频区域之间的联系,以及视觉特征和预测之间的联系。第三,我们提出了一种基于声音的“注意路径”,它利用互补的音频线索来识别重要的视觉区域。我们使用对比学习来进行基于区域的视听实例识别,这是一个中间任务,它得益于音频中的互补线索来提高对象分类和检测的性能。我们表明,我们的方法,更新噪声地面真实情况,并提供间接和注意路径,与单通道预测相比,在AudioSet和VGGSound数据集上大大提高了性能,即使是使用对比学习的预测。我们的方法在AudioSet上达到了最先进的水平,在目标检测任务中的性能优于以前的弱监督检测器,并且我们的声音定位模块在AudioSet和MUSIC上的性能优于几种最先进的方法。
We tackle the problem of learning object detectors in a noisy environment, which is one of the significant challenges for weakly-supervised learning. We use multimodal learning to help localize objects of interest, but unlike other methods, we treat audio as an auxiliary modality that assists to tackle noise in detection from visual regions. First, we use the audio-visual model to generate new "ground-truth" labels for the training set to remove noise between the visual features and noisy supervision. Second, we propose an "indirect path" between audio and class predictions, which combines the link between visual and audio regions, and the link between visual features and predictions. Third, we propose a sound-based "attention path" which uses the benefit of complementary audio cues to identify important visual regions. We use contrastive learning to perform region-based audio-visual instance discrimination, which serves as an intermediate task and benefits from the complementary cues from audio to boost object classification and detection performance. We show that our methods, which update noisy ground truth and provide indirect and attention paths, greatly boosting performance on the AudioSet and VGGSound datasets compared to single-modality predictions, even ones that use contrastive learning. Our method outperforms previous weakly-supervised detectors for the task of object detection by reaching the state-of-art on AudioSet, and our sound localization module performs better than several state-of-art methods on AudioSet and MUSIC.