Self-supervised Attention Model for Weakly Labeled Audio Event Classification

Self-supervised Attention Model for Weakly Labeled Audio Event Classification
复制标题

DOI:
10.23919/eusipco.2019.8902567
复制
发表时间:
2019-08
期刊:
2019 27th European Signal Processing Conference (EUSIPCO)
影响因子:
--
通讯作者:
B. Kim;Shabnam Ghaffarzadegan
B. Kim;Shabnam Ghaffarzadegan
中科院分区:
其他
文献类型:
--
作者:
B. Kim;Shabnam Ghaffarzadegan

文献摘要

被引文献

相似文献

我们描述了一种新的弱标记的音频事件分类方法的基础上的自监督注意力模型。弱标记框架用于消除对昂贵的数据标记过程的需要,并且部署自监督注意力以帮助模型以与先前注意力模型相比更有效的方式区分弱标记音频片段的相关和不相关部分。当强标签可用时,我们还提出了一个高效的强监督注意力模型。该模型也可以作为自监督模型的上限。具有自监督注意训练的模型的性能与使用强标签训练的强监督模型相当。我们表明,我们的自我监督的注意力方法是特别有益的短音频事件。我们实现了8.8%和17.6%的相对平均平均精度的改善,目前国家的最先进的系统SL-DCASE-17和平衡AudioSet。
We describe a novel weakly labeled Audio Event Classification approach based on a self-supervised attention model. The weakly labeled framework is used to eliminate the need for expensive data labeling procedure and self-supervised attention is deployed to help a model distinguish between relevant and irrelevant parts of a weakly labeled audio clip in a more effective manner compared to prior attention models. We also propose a highly effective strongly supervised attention model when strong labels are available. This model also serves as an upper bound for the self-supervised model. The performances of the model with self-supervised attention training are comparable to the strongly supervised one which is trained using strong labels. We show that our self-supervised attention method is especially beneficial for short audio events. We achieve 8.8% and 17.6% relative mean average precision improvements over the current state-of-the-art systems for SL-DCASE-17and balanced AudioSet.