Audio classification with low-rank matrix representation features

Audio classification with low-rank matrix representation features
复制标题

DOI:
10.1145/2542182.2542197
复制
发表时间:
2013-12
期刊:
ACM Transactions on Intelligent Systems and Technology (TIST)
影响因子:
--
通讯作者:
Ziqiang Shi;Jiqing Han;Tieran Zheng
Ziqiang Shi;Jiqing Han;Tieran Zheng
中科院分区:
其他
文献类型:
--
作者:
Ziqiang Shi;Jiqing Han;Tieran Zheng

文献摘要

相似文献

本文提出了一种基于迹范数最小化的音频分类框架。在此框架下,特征提取和分类都是通过求解相应的凸优化问题和迹范数正则化来实现的。在特征提取方面,采用鲁棒主成分分析(Robust PCA)方法,通过最小化核范数和鲁棒1-范数的组合,提取对白色噪声和音频信号严重失真具有鲁棒性的低秩矩阵特征。这些低秩矩阵特征被馈送到线性分类器,其中通过解决类似的迹范数约束问题来学习权重和偏置。对于这种线性分类器,大多数方法都是在批处理模式下找到参数,即权重矩阵和偏差,这使得它在大规模问题中效率低下。在这篇文章中,我们提出了一个并行的在线框架,使用加速邻近梯度方法。该框架在处理速度和内存开销方面具有优势。此外,由于矩阵分类的正则化公式,Lipschitz常数被显式地给出,从而省略了一般近似梯度法的步长估计,节省了这部分计算量.在真实的数据集上进行的笑/非笑和掌声/非掌声分类实验表明,该框架是有效的,并且具有较强的噪声鲁棒性。
In this article, a novel framework based on trace norm minimization for audio classification is proposed. In this framework, both the feature extraction and classification are obtained by solving corresponding convex optimization problem with trace norm regularization. For feature extraction, robust principle component analysis (robust PCA) via minimization a combination of the nuclear norm and the ℓ1-norm is used to extract low-rank matrix features which are robust to white noise and gross corruption for audio signal. These low-rank matrix features are fed to a linear classifier where the weight and bias are learned by solving similar trace norm constrained problems. For this linear classifier, most methods find the parameters, that is the weight matrix and bias in batch-mode, which makes it inefficient for large scale problems. In this article, we propose a parallel online framework using accelerated proximal gradient method. This framework has advantages in processing speed and memory cost. In addition, as a result of the regularization formulation of matrix classification, the Lipschitz constant was given explicitly, and hence the step size estimation of the general proximal gradient method was omitted, and this part of computing burden is saved in our approach. Extensive experiments on real data sets for laugh/non-laugh and applause/non-applause classification indicate that this novel framework is effective and noise robust.