Cross-domain Deep Feature Combination for Bird Species Classification with Audio-visual Data

Cross-domain Deep Feature Combination for Bird Species Classification with Audio-visual Data
复制标题

DOI:
10.1587/transinf.2018edp7383
复制
发表时间:
2018-11
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
B. Naranchimeg;Chao Zhang-;T. Akashi
B. Naranchimeg;Chao Zhang-;T. Akashi
中科院分区:
其他
文献类型:
--
作者:
B. Naranchimeg;Chao Zhang-;T. Akashi

文献摘要

相似文献

近十年来,随着深卷积神经网络(CNN)的发展,许多最先进的图像分类算法和音频分类算法都取得了显著的成就。然而,大多数工作只利用了单一类型的训练数据。在这篇文章中,我们提出了一种利用视觉(图像)和音频(声音)数据相结合的方法来对鸟类进行分类的研究,该方法到目前为止一直是稀疏处理的。具体地说,我们提出了基于CNN的三种融合策略(早、中、晚)的多通道学习模型,以解决训练数据跨域组合的问题。我们提出的方法的优点在于,我们不仅可以利用CNN从图像和音频数据(频谱图)中提取特征,而且可以将不同模式的特征组合在一起。在实验中,我们在一个全面的CUB-200-2011标准数据集上训练和评估网络结构,结合我们最初收集的关于数据种类的音频数据集。我们观察到,使用这两种数据组合的模型比仅使用其中一种类型的数据训练的模型性能更好。我们还表明,迁移学习可以显著提高分类性能。
In recent decade, many state-of-the-art algorithms on image classification as well as audio classification have achieved noticeable successes with the development of deep convolutional neural network (CNN). However, most of the works only exploit single type of training data. In this paper, we present a study on classifying bird species by exploiting the combination of both visual (images) and audio (sounds) data using CNN, which has been sparsely treated so far. Specifically, we propose CNN-based multimodal learning models in three types of fusion strategies (early, middle, late) to settle the issues of combining training data cross domains. The advantage of our proposed method lies on the fact that We can utilize CNN not only to extract features from image and audio data (spectrogram) but also to combine the features across modalities. In the experiment, we train and evaluate the network structure on a comprehensive CUB-200-2011 standard data set combing our originally collected audio data set with respect to the data species. We observe that a model which utilizes the combination of both data outperforms models trained with only an either type of data. We also show that transfer learning can significantly increase the classification performance.