Shift-Invariant Sparse Coding for Audio Classification

Shift-Invariant Sparse Coding for Audio Classification
复制标题

DOI:
--
复制
发表时间:
2007
期刊:
--
影响因子:
--
通讯作者:
Roger B. Grosse;Helen Kwong
Roger B. Grosse;Helen Kwong
中科院分区:
其他
文献类型:
--
作者:
Roger B. Grosse;Helen Kwong

文献摘要

被引文献

相似文献

稀疏编码是一种无监督学习算法,它只在给定未标记数据的情况下学习输入的简洁高级表示;它将每个输入表示为一组基函数的稀疏线性组合。稀疏编码最初应用于人类视觉皮质的建模,也被证明对自学学习有用,其中的目标是解决监督分类任务,因为该任务可以访问来自不同类别的额外的未标记数据,而不是监督学习问题中的数据。移位不变稀疏编码(SISC)是稀疏编码的扩展,它使用所有可能移位中的所有基函数来重建输入(通常是时间序列)。本文提出了一种有效的学习SISC基的算法。我们的方法是基于迭代求解两个大型凸优化问题:第一个问题是计算线性系数的L1正则化线性最小二乘问题,可能有数十万个变量。现有的方法通常使用启发式方法来选择一小部分变量进行优化,但我们提出了一种有效地计算精确解的方法。第二个问题是带约束的线性最小二乘问题,求解基。通过在傅立叶域中对复值变量进行优化,减少了不同变量之间的耦合,从而使问题得到有效的解决。我们表明,SISC对语音和音乐的高级学习表示为这些领域中的分类任务提供了有用的特征。当应用于分类时,在一定条件下,学习的特征优于最先进的光谱和倒谱特征。
Sparse coding is an unsupervised learning algorithm that learns a succinct high-level representation of the inputs given only unlabeled data; it represents each input as a sparse linear combination of a set of basis functions. Originally applied to modeling the human visual cortex, sparse coding has also been shown to be useful for self-taught learning, in which the goal is to solve a supervised classification task given access to additional unlabeled data drawn from different classes than that in the supervised learning problem. Shift-invariant sparse coding (SISC) is an extension of sparse coding which reconstructs a (usually time-series) input using all of the basis functions in all possible shifts. In this paper, we present an efficient algorithm for learning SISC bases. Our method is based on iteratively solving two large convex optimization problems: The first, which computes the linear coefficients, is an L1-regularized linear least squares problem with potentially hundreds of thousands of variables. Existing methods typically use a heuristic to select a small subset of the variables to optimize, but we present a way to efficiently compute the exact solution. The second, which solves for bases, is a constrained linear least squares problem. By optimizing over complex-valued variables in the Fourier domain, we reduce the coupling between the different variables, allowing the problem to be solved efficiently. We show that SISC’s learned high-level representations of speech and music provide useful features for classification tasks within those domains. When applied to classification, under certain conditions the learned features outperform state of the art spectral and cepstral features.