Analyzing the Potential of Pre-Trained Embeddings for Audio Classification Tasks

Analyzing the Potential of Pre-Trained Embeddings for Audio Classification Tasks
复制标题

DOI:
10.23919/eusipco47968.2020.9287743
复制
发表时间:
2021-01
期刊:
2020 28th European Signal Processing Conference (EUSIPCO)
影响因子:
--
通讯作者:
S. Grollmisch;Estefanía Cano;Christian Kehling;Michael Taenzer
S. Grollmisch;Estefanía Cano;Christian Kehling;Michael Taenzer
中科院分区:
其他
文献类型:
--
作者:
S. Grollmisch;Estefanía Cano;Christian Kehling;Michael Taenzer

文献摘要

被引文献

相似文献

在深度学习的背景下,大量训练数据的可用性可以对模型的性能发挥关键作用。最近,几种用于音频分类的模型已经在大型数据集上以监督或自监督的方式进行了预训练,以学习复杂的特征表示,即所谓的嵌入。然后,这些嵌入可以从较小的数据集中提取出来,并用于训练后续的分类器。例如,在音频事件检测(AED)领域中,使用这些特征的分类器已经实现了高准确性,而不需要额外的领域知识。本文评估了三个国家的最先进的嵌入六个音频分类任务,从音乐信息检索和工业声音分析领域。通过分析分类器结构、文件预测的融合方法、训练数据量和嵌入的初始训练域对分类精度的影响,系统地评估了嵌入。为了更好地理解预训练步骤的影响,还将结果与从头开始训练的模型获得的结果进行了比较。平均而言,OpenL3嵌入在线性SVM分类器中表现最好。对于减少的训练示例量,OpenL3的性能优于初始基线。
In the context of deep learning, the availability of large amounts of training data can play a critical role in a model’s performance. Recently, several models for audio classification have been pre-trained in a supervised or self-supervised fashion on large datasets to learn complex feature representations, socalled embeddings. These embeddings can then be extracted from smaller datasets and used to train subsequent classifiers. In the field of audio event detection (AED) for example, classifiers using these features have achieved high accuracy without the need of additional domain knowledge. This paper evaluates three state-of-the-art embeddings on six audio classification tasks from the fields of music information retrieval and industrial sound analysis. The embeddings are systematically evaluated by analyzing the influence on classification accuracy of classifier architecture, fusion methods for file-wise predictions, amount of training data, and initial training domain of the embeddings. To better understand the impact of the pre-training step, results are also compared with those acquired with models trained from scratch. On average, the OpenL3 embeddings performed best with a linear SVM classifier. For a reduced amount of training examples, OpenL3 outperforms the initial baseline.