Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures

Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
复制标题

DOI:
10.1109/tmm.2022.3178591
复制
发表时间:
2021-05
影响因子:
7.3
通讯作者:
Sangwook Park;D. Han;Mounya Elhilali
Sangwook Park;D. Han;Mounya Elhilali
中科院分区:
计算机科学1区
文献类型:
--
作者:
Sangwook Park;D. Han;Mounya Elhilali

文献摘要

被引文献

相似文献

声音事件检测是音频标记的一个重要方面,其目的是识别感兴趣的声音并为连续记录中的每个声音事件定义声音类别和时间边界。随着深度神经网络的进步,声音事件检测系统的性能得到了巨大的改善,尽管代价是昂贵的数据收集和标记工作。事实上,当前最先进的方法采用监督训练方法,该方法利用大量数据样本和对应的标签,以便促进对事件的声音类别和时间戳的识别。作为替代方案,目前的研究提出了一种半监督的方法,用于使用平衡自我训练和交叉训练的学生-教师方案从无监督数据生成伪标签。此外,本文还探讨了从网络预测中提取声音间隔的后处理,以进一步提高声音事件检测性能。所提出的方法进行评估的声音事件检测任务的DCASE 2020的挑战。这些方法在DESED数据库的“验证”和“公共评估”集上的结果与半监督学习中的最先进系统相比显示出显着的改进。
Sound event detection is an important facet of audio tagging that aims to identify sounds of interest and define both the sound category and time boundaries for each sound event in a continuous recording. With advances in deep neural networks, there has been tremendous improvement in the performance of sound event detection systems, although at the expense of costly data collection and labeling efforts. In fact, current state-of-the-art methods employ supervised training methods that leverage large amounts of data samples and corresponding labels in order to facilitate identification of sound category and time stamps of events. As an alternative, the current study proposes a semi-supervised method for generating pseudo-labels from unsupervised data using a student-teacher scheme that balances self-training and cross-training. Additionally, this paper explores post-processing which extracts sound intervals from network prediction, for further improvement in sound event detection performance. The proposed approach is evaluated on sound event detection task for the DCASE2020 challenge. The results of these methods on both “validation” and “public evaluation” sets of DESED database show significant improvement compared to the state-of-the art systems in semi-supervised learning.