Pyramidal Temporal Pooling With Discriminative Mapping for Audio Classification

Pyramidal Temporal Pooling With Discriminative Mapping for Audio Classification
复制标题

用于音频分类的具有判别性映射的金字塔时间池

DOI:
10.1109/taslp.2020.2966868
复制
发表时间:
2020-01
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Jiqing Han
Jiqing Han
中科院分区:
其他
文献类型:
--
作者:
Liwen Zhang;Ziqiang Shi;Jiqing Han

文献摘要

参考文献

被引文献

相似文献

音频信号是时间结构的数据,学习其包含时间信息的区分表示是音频分类的关键。在本文中,我们提出了一种具有层次金字塔结构的音频表示学习方法,称为金字塔时间池(PTP),旨在捕获整个音频样本的时间信息。通过在多个局部时间池层上叠加一个全局时间池层,PTP能够以一种无监督的方式捕捉输入特征序列的高层时间动态。此外,在顶级全局时间池层,我们联合优化了可学习判别映射(DM)和Softmax分类器。为此,还提出了一种用于区分音频表示和分类器的联合学习方法DM-PTP。DM-PTP将时间编码作为一个双层优化问题的底层约束,在保持整个序列的时间信息的同时,能够产生区分表示。对于任意持续时间的音频样本,我们的PTP和DM-PTP都可以将任意长度的输入特征序列编码成固定长度的表示。在不使用任何数据增强和集成学习方法的情况下,PTP和DM-PTP在音频事件识别(AER)数据集上的性能都优于最先进的CNN,并且在DCASE 2018声学场景分类(ASC)数据集上的性能可以与挑战中的其他最佳模型相媲美。
Audio signals are temporally-structured data, and learning their discriminative representations containing temporal information is crucial for the audio classification. In this article, we propose an audio representation learning method with a hierarchical pyramid structure called pyramidal temporal pooling (PTP) which aims to capture the temporal information of an entire audio sample. By stacking a global temporal pooling layer on multiple local temporal pooling layers, the PTP can capture the high-level temporal dynamics of the input feature sequence in an unsupervised way. Furthermore, in the top global temporal pooling layer, we jointly optimize a learnable discriminative mapping (DM) and a softmax classifier. Such that, a joint learning method for the discriminative audio representations and the classifier called DM-PTP is also presented. By treating the temporal encoding as a low-level constraint of a bi-level optimization problem, the DM-PTP can produce the discriminative representation while maintaining the temporal information of the whole sequence. For an audio sample with an arbitrary time duration, both our PTP and DM-PTP can encode the input feature sequence with arbitrary length into a fixed-length representation. Without using any data augmentation and ensemble learning methods, both PTP and DM-PTP outperform the state-of-the-art CNNs on the audio event recognition (AER) dataset, and can achieve comparable performance on the DCASE 2018 acoustic scene classification (ASC) dataset compared with other best models in the challenge.
DOI: --
发表时间: 2017-09
期刊: ArXiv
影响因子: --
作者:
Matthias Meyer;Lukas Cavigelli;L. Thiele
通讯作者: Matthias Meyer;Lukas Cavigelli;L. Thiele
DOI: 10.1007/978-3-540-68585-2_32
发表时间: 2007
期刊: --
影响因子: --
作者:
C. Zieger
通讯作者: C. Zieger
DOI: --
发表时间: 2018-07
期刊: --
影响因子: --
作者:
A. Mesaros;Toni Heittola;T. Virtanen
通讯作者: A. Mesaros;Toni Heittola;T. Virtanen
DOI: 10.1109/icassp.2018.8461975
发表时间: 2017-10
期刊: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
Yong Xu;Qiuqiang Kong;Wenwu Wang;Mark D. Plumbley
通讯作者: Yong Xu;Qiuqiang Kong;Wenwu Wang;Mark D. Plumbley
DOI: 10.1109/taslp.2014.2375575
发表时间: 2015
期刊: IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子: --
作者:
A. Rakotomamonjy;G. Gasso
通讯作者: A. Rakotomamonjy;G. Gasso