Robust sound event classification with bilinear multi-column ELM-AE and two-stage ensemble learning

Robust sound event classification with bilinear multi-column ELM-AE and two-stage ensemble learning
复制标题

使用双线性多列 ELM-AE 和两阶段集成学习进行鲁棒声音事件分类

DOI:
10.1186/s13636-017-0109-1
复制
发表时间:
2017
影响因子:
2.4
通讯作者:
Li Yan
Li Yan
中科院分区:
计算机科学4区
文献类型:
--
作者:
Zhang Junjie;Yin Jie;Zhang Qi;Shi Jun;Li Yan

文献摘要

相似文献

近年来,声音事件自动分类(SEC)受到越来越多的关注。特征提取是SEC系统中的一个关键因素,深度神经网络(DNN)算法已经达到了SEC的最先进性能。基于极端学习机器的自动编码器(ELM-AE)是一种新的深度学习算法,它具有出色的表示性能和非常快速的训练过程。然而,ELM-AE遭受不稳定的问题。在这项工作中,提出了一种双线性多列ELM-AE(B-MC-ELM-AE)算法,以改善原始ELM-AE的鲁棒性,稳定性和特征表示,然后将其应用于学习声音信号的特征表示。此外,B-MC-ELM-AE和两阶段集成学习(TsEL)为基础的特征学习和分类框架,然后开发执行强大的和有效的SEC。在真实的世界计算伙伴关系声音场景数据库上的实验结果表明,所提出的SEC框架优于最先进的DNN算法。
The automatic sound event classification (SEC) has attracted a growing attention in recent years. Feature extraction is a critical factor in SEC system, and the deep neural network (DNN) algorithms have achieved the state-of-the-art performance for SEC. The extreme learning machine-based auto-encoder (ELM-AE) is a new deep learning algorithm, which has both an excellent representation performance and very fast training procedure. However, ELM-AE suffers from the problem of unstability. In this work, a bilinear multi-column ELM-AE (B-MC-ELM-AE) algorithm is proposed to improve the robustness, stability, and feature representation of the original ELM-AE, which is then applied to learn feature representation of sound signals. Moreover, a B-MC-ELM-AE and two-stage ensemble learning (TsEL)-based feature learning and classification framework is then developed to perform the robust and effective SEC. The experimental results on the Real World Computing Partnership Sound Scene Database show that the proposed SEC framework outperforms the state-of-the-art DNN algorithm.