Multi-Task Learning for Acoustic Event Detection Using Event and Frame Position Information

Multi-Task Learning for Acoustic Event Detection Using Event and Frame Position Information
复制标题

DOI:
10.1109/tmm.2019.2933330
复制
发表时间:
2020-03
影响因子:
7.3
通讯作者:
Xianjun Xia;R. Togneri;Ferdous Sohel;Yuanjun Zhao;David Huang
Xianjun Xia;R. Togneri;Ferdous Sohel;Yuanjun Zhao;David Huang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Xianjun Xia;R. Togneri;Ferdous Sohel;Yuanjun Zhao;David Huang

文献摘要

被引文献

相似文献

声事件检测对声信号进行处理,以确定声音类型并估计音频事件边界。基于多标签分类的方法通常用于检测帧方式的事件类型,并应用中值滤波来确定发生的声学事件。然而,多标签分类器仅针对声音事件类型进行训练,而忽略了音频事件中的帧位置。针对这一问题,本文提出构建一个基于联合学习的多任务系统。第一任务执行声事件类型检测,第二任务是预测帧位置信息。通过在两个任务之间共享表示,我们可以通过将各自的噪声模式平均为隐式正则化来使声学模型比原始分类器具有更好的泛化能力。在单声道UPC-TALP和多声道TUT声音事件数据集上的实验结果表明,与基准AED系统相比,该联合学习方法获得了更低的错误率和更高的F-Score。
Acoustic event detection deals with the acoustic signals to determine the sound type and to estimate the audio event boundaries. Multi-label classification based approaches are commonly used to detect the frame wise event types with a median filter applied to determine the happening acoustic events. However, the multi-label classifiers are trained only on the acoustic event types ignoring the frame position within the audio events. To deal with this, this paper proposes to construct a joint learning based multi-task system. The first task performs the acoustic event type detection and the second task is to predict the frame position information. By sharing representations between the two tasks, we can enable the acoustic models to generalize better than the original classifier by averaging respective noise patterns to be implicitly regularized. Experimental results on the monophonic UPC-TALP and the polyphonic TUT Sound Event datasets demonstrate the superior performance of the joint learning method by achieving lower error rate and higher F-score compared to the baseline AED system.