Adaptive Pooling Operators for Weakly Labeled Sound Event Detection

Adaptive Pooling Operators for Weakly Labeled Sound Event Detection
复制标题

DOI:
10.1109/taslp.2018.2858559
复制
发表时间:
2018-04
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Brian McFee;J. Salamon;J. Bello
Brian McFee;J. Salamon;J. Bello
中科院分区:
其他
文献类型:
--
作者:
Brian McFee;J. Salamon;J. Bello

文献摘要

被引文献

相似文献

声音事件检测(SED)方法的任务是通过活动声源的存在来标记音频记录的片段。SED通常被视为一个有监督的机器学习问题,需要在录音中的每个时刻对每个声源的存在或不存在进行强注释。然而,这种类型的强注释对于人类注释者来说是劳动密集型和成本密集型的,这限制了SED方法的实际可扩展性。在本文中,我们将SED视为一个多实例学习(MIL)问题,其中训练标签在短摘录上是静态的,表明声源的存在或不存在,但不是它们的时间局部性。然而,这些模型仍然必须产生时间动态预测,在训练期间与静态标签进行比较时,这些预测必须被聚合(合并)。为了促进这种聚合,我们开发了一个家庭的自适应池运营商,称为自动池,它顺利插入常见的池运营商,如最小,最大,或平均池,并自动适应的声源的特性的问题。我们在三个数据集上评估了所提出的池化算子,并证明了在每种情况下,所提出的方法在静态预测方面都优于非自适应池化算子,并且几乎与使用强大的动态注释训练的模型的性能相匹配。所提出的方法结合卷积神经网络进行评估,但可以很容易地应用于任何可微模型的时间序列标签预测。虽然本文侧重于SED应用,所提出的方法是通用的,可以广泛应用于任何领域的MIL问题。
Sound event detection (SED) methods are tasked with labeling segments of audio recordings by the presence of active sound sources. SED is typically posed as a supervised machine learning problem, requiring strong annotations for the presence or absence of each sound source at every time instant within the recording. However, strong annotations of this type are both labor- and cost-intensive for human annotators to produce, which limits the practical scalability of SED methods. In this paper, we treat SED as a multiple instance learning (MIL) problem, where training labels are static over a short excerpt, indicating the presence or absence of sound sources but not their temporal locality. The models, however, must still produce temporally dynamic predictions, which must be aggregated (pooled) when comparing against static labels during training. To facilitate this aggregation, we develop a family of adaptive pooling operators—referred to as autopool—which smoothly interpolate between common pooling operators, such as min-, max-, or average-pooling, and automatically adapt to the characteristics of the sound sources in question. We evaluate the proposed pooling operators on three datasets, and demonstrate that in each case, the proposed methods outperform nonadaptive pooling operators for static prediction, and nearly match the performance of models trained with strong, dynamic annotations. The proposed method is evaluated in conjunction with convolutional neural networks, but can be readily applied to any differentiable model for time-series label prediction. While this paper focuses on SED applications, the proposed methods are general, and could be applied widely to MIL problems in any domain.