Improving Semi-Supervised Learning for Audio Classification with FixMatch

Improving Semi-Supervised Learning for Audio Classification with FixMatch
复制标题

DOI:
10.3390/electronics10151807
复制
发表时间:
2021-08-01
期刊:
影响因子:
2.9
通讯作者:
Cano, Estefania
Cano, Estefania
中科院分区:
工程技术3区
文献类型:
--
作者:
Grollmisch, Sascha;Cano, Estefania

文献摘要

被引文献

相似文献

使用半监督学习 (SSL) 在神经网络的训练过程中包含未标记数据已在图像领域显示出令人印象深刻的结果,其中仅使用一小部分标记数据就获得了最先进的结果。最近的 SSL 方法之间的共同点是它们强烈依赖于未注释数据的增强。对于音频数据来说,这还没有被探索过。在这项工作中,使用最先进的 FixMatch 方法的 SSL 在三个音频分类任务上进行了评估,包括音乐、工业声音和声学场景。 FixMatch 的性能与使用 Mean Teacher 方法从头开始训练的卷积神经网络 (CNN)、迁移学习和 SSL 进行了比较。此外,还介绍了一种简单而有效的方法,用于为 FixMatch 选择合适的增强方法。采用建议的修改后的 FixMatch 始终优于 Mean Teacher 和从头开始训练的 CNN。对于工业声音和音乐数据集,使用完整数据集的 CNN 基线性能只需不到 5% 的初始训练数据即可达到,这证明了最新 SSL 方法在音频数据方面的潜力。仅在声学场景分类中最具挑战性的数据集上,迁移学习的表现才优于 FixMatch,这表明仍有改进的空间。
Including unlabeled data in the training process of neural networks using Semi-Supervised Learning (SSL) has shown impressive results in the image domain, where state-of-the-art results were obtained with only a fraction of the labeled data. The commonality between recent SSL methods is that they strongly rely on the augmentation of unannotated data. This is vastly unexplored for audio data. In this work, SSL using the state-of-the-art FixMatch approach is evaluated on three audio classification tasks, including music, industrial sounds, and acoustic scenes. The performance of FixMatch is compared to Convolutional Neural Networks (CNN) trained from scratch, Transfer Learning, and SSL using the Mean Teacher approach. Additionally, a simple yet effective approach for selecting suitable augmentation methods for FixMatch is introduced. FixMatch with the proposed modifications always outperformed Mean Teacher and the CNNs trained from scratch. For the industrial sounds and music datasets, the CNN baseline performance using the full dataset was reached with less than 5% of the initial training data, demonstrating the potential of recent SSL methods for audio data. Transfer Learning outperformed FixMatch only for the most challenging dataset from acoustic scene classification, showing that there is still room for improvement.