Polyphonic Sound Event and Sound Activity Detection: A Multi-Task Approach

Polyphonic Sound Event and Sound Activity Detection: A Multi-Task Approach
复制标题

和弦声音事件和声音活动检测:多任务方法

DOI:
10.1109/waspaa.2019.8937193
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Pankajakshan A
Pankajakshan A
中科院分区:
--
文献类型:
--
作者:
Pankajakshan A

文献摘要

参考文献

被引文献

相似文献

由于声音事件的动态复调电平、强度和持续时间,现实录音中的复调声音事件检测 (SED) 是一项具有挑战性的任务。当前的和弦 SED 系统无法明确地对声音事件的时间结构进行建模,而是尝试查看每个音频帧中存在哪些声音事件。因此,事件级检测性能远低于分段检测性能。在这项工作中,我们提出了一种联合模型方法,使用多任务学习设置来改进声音事件的时间定位。第一个任务预测每个时间帧出现哪些声音事件;我们将此分支称为“声音事件检测(SED)模型”,而第二个任务则预测每一帧是否存在声音事件;我们将此分支称为“声音活动检测(SAD)模型”。我们通过将单个任务预测聚合在一起的两个任务的单独实现进行比较来验证所提出的联合模型。我们在 URBAN-SED 数据集上的实验表明,所提出的联合模型可以减轻假阳性(FP)和假阴性(FN)错误,并改善分段和事件方面的指标。
Polyphonic Sound Event Detection (SED) in real-world recordings is a challenging task because of the dynamic polyphony level, intensity, and duration of sound events. Current polyphonic SED systems fail to model the temporal structure of sound events explicitly and instead attempt to look at which sound events are present at each audio frame. Consequently, the event-wise detection performance is much lower than the segment-wise detection performance. In this work, we propose a joint model approach to improve the temporal localization of sound events using a multi-task learning setup. The first task predicts which sound events are present at each time frame; we call this branch `Sound Event Detection (SED) model', while the second task predicts if a sound event is present or not at each frame; we call this branch `Sound Activity Detection (SAD) model'. We verify the proposed joint model by comparing it with a separate implementation of both tasks aggregated together from individual task predictions. Our experiments on the URBAN-SED dataset show that the proposed joint model can alleviate False Positive (FP) and False Negative (FN) errors and improve both the segment-wise and the event-wise metrics.
现实音频中的声音事件检测
DOI: --
发表时间: 2016
期刊:
影响因子: --
作者:
Dmitrii Ubskii;A. Pugachev
通讯作者: A. Pugachev
DOI: 10.1109/taslp.2017.2690576
发表时间: 2017-06
期刊: IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子: --
作者:
Emmanouil Benetos;G. Lafay;M. Lagrange;Mark D. Plumbley
通讯作者: Emmanouil Benetos;G. Lafay;M. Lagrange;Mark D. Plumbley
DOI: 10.1109/waspaa.2017.8170052
发表时间: 2017-10
期刊: 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
影响因子: --
作者:
J. Salamon;D. MacConnell;M. Cartwright;P. Li;J. Bello
通讯作者: J. Salamon;D. MacConnell;M. Cartwright;P. Li;J. Bello
能够使用深度学习从面部照片中区分阿尔茨海默病
DOI: --
发表时间: 2020
期刊:
影响因子: --
作者:
Manabu Kokubo;Akihiro Hirashiki;Takahiro Kamihara;Atsuya Shimizu;Hidenori Arai;亀山祐美,亀山征史,深澤誠,飯塚友道,飯島勝矢,田中友規,矢可部満隆,小島太郎,小川純人,秋下雅弘
通讯作者: 亀山祐美,亀山征史,深澤誠,飯塚友道,飯島勝矢,田中友規,矢可部満隆,小島太郎,小川純人,秋下雅弘