Automatic large-scale classification of bird sounds is strongly improved by unsupervised feature learning.

Automatic large-scale classification of bird sounds is strongly improved by unsupervised feature learning.
复制标题

DOI:
10.7717/peerj.488
复制
发表时间:
2014
期刊:
影响因子:
2.7
通讯作者:
Plumbley MD
Plumbley MD
中科院分区:
生物学3区
文献类型:
--
作者:
Stowell D;Plumbley MD

文献摘要

参考文献

被引文献

相似文献

根据鸟类的声音进行鸟类自动分类是一种在生态学、保护监测和声音交流研究中越来越重要的计算工具。为了使分类在实践中有用,关键是要提高其准确性,同时确保它可以在大数据规模下运行。许多方法使用基于频谱图类型数据的声学测量,例如Mel频率倒谱系数(MFCC)特征,其表示频谱信息的手动设计的概要。然而,机器学习领域最近的研究表明,从数据中自动学习的特征往往优于手动设计的特征变换。特征学习可以在大规模和“无监督”的情况下进行,这意味着它不需要手动标记数据,但它可以提高分类等“监督”任务的性能。在这项工作中,我们介绍了一种技术,从大量的鸟的声音记录,已被证明在其他领域有用的技术的启发,功能学习。我们实验比较十二个不同的功能表示来自梅尔频谱(其中六个使用这种技术),使用四个大型和不同的数据库的鸟类发声,分类使用随机森林分类器。我们证明,在我们的分类任务中,MFCC往往会导致更差的性能比原始梅尔光谱数据,他们来自。相反,我们证明了无监督特征学习在模型训练后,在不增加计算复杂性的情况下,大大提高了MFCC和Mel谱。这种提升对于大规模的单标签分类任务尤其明显。通过我们的程序学习的频谱-时间激活类似于从鸟类初级听觉前脑计算的频谱-时间感受野。然而,对于我们的一个数据集,它包含大量的音频数据,但很少注释,提高性能是不可察觉的。通过进一步的实证分析,我们研究了数据集特征和特征表示选择之间的相互作用。
Automatic species classification of birds from their sound is a computational tool of increasing importance in ecology, conservation monitoring and vocal communication studies. To make classification useful in practice, it is crucial to improve its accuracy while ensuring that it can run at big data scales. Many approaches use acoustic measures based on spectrogram-type data, such as the Mel-frequency cepstral coefficient (MFCC) features which represent a manually-designed summary of spectral information. However, recent work in machine learning has demonstrated that features learnt automatically from data can often outperform manually-designed feature transforms. Feature learning can be performed at large scale and “unsupervised”, meaning it requires no manual data labelling, yet it can improve performance on “supervised” tasks such as classification. In this work we introduce a technique for feature learning from large volumes of bird sound recordings, inspired by techniques that have proven useful in other domains. We experimentally compare twelve different feature representations derived from the Mel spectrum (of which six use this technique), using four large and diverse databases of bird vocalisations, classified using a random forest classifier. We demonstrate that in our classification tasks, MFCCs can often lead to worse performance than the raw Mel spectral data from which they are derived. Conversely, we demonstrate that unsupervised feature learning provides a substantial boost over MFCCs and Mel spectra without adding computational complexity after the model has been trained. The boost is particularly notable for single-label classification tasks at large scale. The spectro-temporal activations learned through our procedure resemble spectro-temporal receptive fields calculated from avian primary auditory forebrain. However, for one of our datasets, which contains substantial audio data but few annotations, increased performance is not discernible. We study the interaction between dataset characteristics and choice of feature representation through further empirical analysis.
DOI: 10.1109/tassp.1980.1163420
发表时间: 1980-01-01
期刊: IEEE TRANSACTIONS ON ACOUSTICS SPEECH AND SIGNAL PROCESSING
影响因子: --
作者:
DAVIS, SB;MERMELSTEIN, P
通讯作者: MERMELSTEIN, P
DOI: 10.1121/1.415968
发表时间: 1996-08-01
影响因子: 2.4
作者:
Anderson, SE;Dave, AS;Margoliash, D
通讯作者: Margoliash, D
DOI: 10.7717/peerj.103
发表时间: 2013
期刊: PeerJ
影响因子: 2.7
作者:
Aide TM;Corrada-Bravo C;Campos-Cerqueira M;Milan C;Vega G;Alvarez R
通讯作者: Alvarez R
DOI: 10.1111/2041-210x.12060
发表时间: 2013-07-01
影响因子: 6.6
作者:
Digby, Andrew;Towsey, Michael;Teal, Paul D.
通讯作者: Teal, Paul D.
DOI: 10.1016/j.ecoinf.2009.06.005
发表时间: 2009-09-01
影响因子: 5.1
作者:
Acevedo, Miguel A.;Corrada-Bravo, Carlos J.;Aide, T. Mitchell
通讯作者: Aide, T. Mitchell