Fusing shallow and deep learning for bioacoustic bird species classification

Fusing shallow and deep learning for bioacoustic bird species classification
复制标题

DOI:
10.1109/icassp.2017.7952134
复制
发表时间:
2017-03
期刊:
2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
J. Salamon;J. Bello;Andrew Farnsworth;S. Kelling
J. Salamon;J. Bello;Andrew Farnsworth;S. Kelling
中科院分区:
其他
文献类型:
--
作者:
J. Salamon;J. Bello;Andrew Farnsworth;S. Kelling

文献摘要

被引文献

相似文献

根据生物的发声将其自动分类为物种,将极大地促进监测生物多样性的能力,在生态学领域有着广泛的应用。特别是,候鸟飞行叫声的自动分类可以为迁徙过程中发声的鸟类提供新的生物学见解和保护应用。在本文中,我们探讨国家的最先进的分类技术,大词汇量的鸟类物种分类飞行呼叫。特别是,我们将基于无监督字典学习的“浅层学习”方法与结合数据增强的深度卷积神经网络进行了对比。我们发现,这两个模型执行的数据集上的5428飞行呼叫跨越43个不同的物种,都显着优于MFCC基线。最后,我们表明,通过使用简单的后期融合方法组合模型,我们可以进一步改善结果,获得最先进的分类精度为0.96。
Automated classification of organisms to species based on their vocalizations would contribute tremendously to abilities to monitor biodiversity, with a wide range of applications in the field of ecology. In particular, automated classification of migrating birds' flight calls could yield new biological insights and conservation applications for birds that vocalize during migration. In this paper we explore state-of-the-art classification techniques for large-vocabulary bird species classification from flight calls. In particular, we contrast a “shallow learning” approach based on unsupervised dictionary learning with a deep convolutional neural network combined with data augmentation. We show that the two models perform comparably on a dataset of 5428 flight calls spanning 43 different species, with both significantly outperforming an MFCC baseline. Finally, we show that by combining the models using a simple late-fusion approach we can further improve the results, obtaining a state-of-the-art classification accuracy of 0.96.