Challenges in Applying Audio Classification Models to Datasets Containing Crucial Biodiversity Information

Challenges in Applying Audio Classification Models to Datasets Containing Crucial Biodiversity Information
复制标题

将音频分类模型应用于包含重要生物多样性信息的数据集的挑战

DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
C. Schurgers
C. Schurgers
中科院分区:
--
文献类型:
--
作者:
Jacob I. Ayers;Yaman Jandali;Yoo;Gabriel Steinberg;Erika Joun;M. Tobler;Ian Ingram;R. Kastner;C. Schurgers

文献摘要

参考文献

被引文献

相似文献

自然声景的声学特征可以揭示气候变化对生物多样性的影响。硬件成本、人力时间和专门用于标记音频的专业知识是在生态系统的代表性部分进行声学调查的障碍。随着低成本、易于使用的开源硬件的出现,以及机器学习领域的扩展,这些障碍正在迅速消失,机器学习领域提供了预先训练的神经网络,可以对检索的声学数据进行测试。被动声学监测(PAM)面临的一个持续挑战是,神经网络对现场收集的音频记录缺乏可靠性,这些音频记录中包含重要的生物多样性信息,否则公开可用的训练和测试集会显示出有希望的结果。为了证明这一挑战,我们测试了一种混合循环神经网络(RNN)和卷积神经网络(CNN)二元分类器,在两台秘鲁鸟类音频集上训练鸟类存在/不存在。在Xeno-canto和b谷歌的AudioSet本体收集的数据集上,RNN在接收者操作特征(AUROC)下的面积达到95%,而在秘鲁亚马逊地区Madre de Dios地区收集的现场录音分层随机样本中,这一比例为65%。为了缓解这种差异,我们在网络的训练过程中应用了各种音频数据增强技术,从而使整个现场录音的AUROC达到77%。
The acoustic signature of a natural soundscape can reveal consequences of climate change on biodiversity. Hardware costs, human labor time, and expertise dedicated to labeling audio are impediments to conducting acoustic surveys across a representative portion of an ecosystem. These barriers are quickly eroding away with the advent of low-cost, easy to use, open source hardware and the expansion of the machine learning field providing pre-trained neural networks to test on retrieved acoustic data. One consistent challenge in passive acoustic monitoring (PAM) is a lack of reliability from neural networks on audio recordings collected in the field that contain crucial biodiversity information that otherwise show promising results from publicly available training and test sets. To demonstrate this challenge, we tested a hybrid recurrent neural network (RNN) and convolutional neural network (CNN) binary classifier trained for bird presence/absence on two Peruvian bird audiosets. The RNN achieved an area under the receiver operating characteristics (AUROC) of 95% on a dataset collected from Xeno-canto and Google’s AudioSet ontology in contrast to 65% across a stratified random sample of field recordings collected from the Madre de Dios region of the Peruvian Amazon. In an attempt to alleviate this discrepancy, we applied various audio data augmentation techniques in the network’s training process which led to an AUROC of 77% across the field recordings.
DOI: 10.1111/2041-210x.12955
发表时间: 2018-05-01
影响因子: 6.6
作者:
Hill, Andrew P.;Prince, Peter;Rogers, Alex
通讯作者: Rogers, Alex