Self-taught learning

Self-taught learning
复制标题

DOI:
--
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
A. Ng;Rajat Raina
A. Ng;Rajat Raina
中科院分区:
其他
文献类型:
--
作者:
A. Ng;Rajat Raina

文献摘要

被引文献

相似文献

我们介绍了一种新的机器学习框架,称为自学学习,用于在监督分类任务中使用未标记的数据。该框架不要求未标记的数据遵循受监督任务的类别标签,或来自相同的生成性分布。这种未标记的数据通常比以前研究的框架(如半监督学习)更容易获得。在本文中,我们证明了自学学习可以成功地应用于各种困难的机器学习问题。我们工作的中心是一种基于名为“稀疏编码”的优化问题的自学学习算法。该算法使用未标记的数据来学习复杂的高维输入的新表示,然后在该表示上应用监督学习。该表示法捕获了输入的更高级别的方面,并显著提高了许多测试域的分类性能,包括计算机视觉、音频识别和文本分类。我们为该模型的平移不变版本提出了高效的稀疏编码算法,该算法可以应用于音频和图像数据。我们还将该模型推广到更广泛的输入类别,包括用以前的算法难以处理的领域,并将该模型应用于文本分类和机器人感知任务。综上所述,这些实验表明,使用自学学习框架,机器学习可以应用于比以前更难的问题。这些自学的学习算法在被允许使用大量未标记的数据(数百万个示例)学习丰富的模型(具有数百万个参数)时工作得最好。不幸的是,按照目前的方法,学习如此丰富的模型可能需要数周时间。此外,这些方法需要快速、顺序的更新,并且使用当前的算法,不利于在分布式集群上并行化。为了将自学学习应用于如此大规模的问题,我们展示了图形处理器硬件(在大多数现代台式机中可用)可以用于大规模并行化算法。使用一种新的内在并行算法,稀疏编码算法可以很容易地在图形处理器上实现,并且我们证明这可以将学习时间从大约三周减少到一天。最后,我们考虑了使用未标记数据学习分层表示的自学学习方法。我们提出了利用图形处理器对这类层次模型进行无监督学习的一般原则,并证明了针对流行的深度信念网络模型的慢学习算法可以成功地并行化。这种实施比优化的CPU实施快70倍,将学习时间从几周减少到几个小时,并代表了学习大型深度信念网络的最先进水平。
We introduce a new machine learning framework called self-taught learning for using unlabeled data in supervised classification tasks. This framework does not require that the unlabeled data follow the class labels of the supervised task, or arise from the same generative distribution. Such unlabeled data is often significantly easier to obtain than in previously studied frameworks such as semi-supervised learning. In this thesis, we demonstrate that self-taught learning can be applied successfully to a variety of hard machine learning problems. The centerpiece of our work is a self-taught learning algorithm based on an optimization problem called "sparse coding." This algorithm uses unlabeled data to learn a new representation for complex, high-dimensional inputs, and then applies supervised learning over this representation. The representation captures higher-level aspects of the input, and significantly improves classification performance on many test domains, including computer vision, audio recognition and text classification. We present efficient sparse coding algorithms for a translation-invariant version of the model, that can be applied to audio and image data. We also generalize the model to a much broader class of inputs, including domains that are hard to handle with previous algorithms, and apply the model to text classification and a robotic perception task. Taken together, these experiments demonstrate that using the self-taught learning framework, machine learning can be applied to much harder problems than previously possible. These self-taught learning algorithms work best when they are allowed to learn rich models (with millions of parameters) using large amounts of unlabeled data (millions of examples). Unfortunately, with current methods, it can take weeks to learn such rich models. Further, these methods require fast, sequential updates, and with current algorithms, are not conducive to being parallelized on a distributed cluster. To apply self-taught learning to such large-scale problems, we show that graphics processor hardware (available in most modern desktops) can be used to massively parallelize the algorithms. Using a new inherently parallel algorithm, the sparse coding algorithm can be easily implemented on graphics processors, and we show that this can reduce the learning time from about three weeks to a single day. Finally, we consider self-taught learning methods that learn hierarchical representations using unlabeled data. We develop general principles for unsupervised learning of such hierarchical models using graphics processors, and show that the slow learning algorithms for the popular deep belief network model can be successfully parallelized. This implementation is up to 70 times faster than an optimized CPU implementation, reduces the learning time from weeks to hours, and represents the state-of-the-art in learning large deep belief networks.