Towards Duration Robust Weakly Supervised Sound Event Detection

Towards Duration Robust Weakly Supervised Sound Event Detection
复制标题

走向持续时间鲁棒弱监督声音事件检测

DOI:
10.1109/taslp.2021.3054313
复制
发表时间:
2021-01-01
影响因子:
5.4
通讯作者:
Yu, Kai
Yu, Kai
中科院分区:
计算机科学2区
文献类型:
--
作者:
Dinkel, Heinrich;Wu, Mengyue;Yu, Kai

文献摘要

被引文献

相似文献

声音事件检测 (SED) 的任务是标记给定音频剪辑中音频事件的缺失或存在及其相应的间隔。虽然 SED 可以使用监督机器学习来完成,其中训练数据完全标记为可以访问每个事件的时间戳和持续时间,但我们的工作重点是弱监督声音事件检测 (WSSED),其中无法获得有关事件持续时间的先验知识。该领域最近的研究重点是提高有关特定评估指标的特定数据集的分段和事件级本地化性能。具体来说,性能良好的事件级本地化需要完全标记的开发子集来获取事件持续时间估计,这显着提高了本地化性能。此外,性能良好的分段级本地化模型以粗尺度(例如 1 秒)输出预测,阻碍了它们在包含非常短事件($< 1 秒)的数据集上的部署。这项工作提出了一个持续时间稳健的 CRNN (CDur) 框架,旨在在分段和事件级本地化方面实现有竞争力的性能。本文提出了一种名为“三重阈值”的新后处理策略,并研究了 WSSED 范围内的两种数据增强方法以及标签平滑方法。我们的模型评估是在 DCASE2017 和 2018 Task 4 数据集以及 URBAN-SED 上完成的。我们的模型在 DCASE2018 和 URBAN-SED 数据集上的性能优于其他方法,而无需事先了解持续时间。特别是,我们的模型能够与 URBAN-SED 数据集上的强标记监督模型具有相似的性能。最后,消融实验表明,在没有后处理的情况下,我们的模型的定位性能下降明显低于其他方法。
Sound event detection (SED) is the task of tagging the absence or presence of audio events and their corresponding interval within a given audio clip. While SED can be done using supervised machine learning, where training data is fully labeled with access to per event timestamps and duration, our work focuses on weakly-supervised sound event detection (WSSED), where prior knowledge about an event's duration is unavailable. Recent research within the field focuses on improving segment- and event-level localization performance for specific datasets regarding specific evaluation metrics. Specifically, well-performing event-level localization requires fully labeled development subsets to obtain event duration estimates, which significantly benefits localization performance. Moreover, well-performing segment-level localization models output predictions at a coarse-scale (e.g., 1 second), hindering their deployment on datasets containing very short events ($< $ 1 second). This work proposes a duration robust CRNN (CDur) framework, which aims to achieve competitive performance in terms of segment- and event-level localization. This paper proposes a new post-processing strategy named “Triple Threshold” and investigates two data augmentation methods along with a label smoothing method within the scope of WSSED. Evaluation of our model is done on the DCASE2017 and 2018 Task 4 datasets, and URBAN-SED. Our model outperforms other approaches on the DCASE2018 and URBAN-SED datasets without requiring prior duration knowledge. In particular, our model is capable of similar performance to strongly-labeled supervised models on the URBAN-SED dataset. Lastly, ablation experiments to reveal that without post-processing, our model's localization performance drop is significantly lower compared with other approaches.