Learning Multiple Layers of Features from Tiny Images

Learning Multiple Layers of Features from Tiny Images
复制标题

DOI:
--
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
A. Krizhevsky
A. Krizhevsky
中科院分区:
其他
文献类型:
--
作者:
A. Krizhevsky

文献摘要

被引文献

相似文献

麻省理工学院和纽约大学的小组从网络中收集了数百万个微小颜色图像的数据集。原则上,这是一个出色的数据集,用于无监督的深层生成模型,但是以前尝试过的研究人员已经发现,从图像中学习了一系列好泡沫。我们展示了如何训练多层生成模型,该模型学会提取有意义的特征,这些特征类似于人类视觉皮层中的特征。使用一种新颖的并行化算法在网络上连接的多个机器之间分配工作,我们展示了如何在合理的时间内进行训练。小型图像数据集的第二个有问题的方面是,没有可靠的类标签,因此很难用于对象识别实验。我们创建了两组可靠的标签。 CIFAR-10集有10个类别的6000个示例,CIFAR-100集有600个非重叠类中的600个示例。使用这些标签,我们表明,通过预先培训一层特征在一组未标记的微型图像上,可以显着改善对象识别。
Groups at MIT and NYU have collected a dataset of millions of tiny colour images from the web. It is, in principle, an excellent dataset for unsupervised training of deep generative models, but previous researchers who have tried this have found it dicult to learn a good set of lters from the images. We show how to train a multi-layer generative model that learns to extract meaningful features which resemble those found in the human visual cortex. Using a novel parallelization algorithm to distribute the work among multiple machines connected on a network, we show how training such a model can be done in reasonable time. A second problematic aspect of the tiny images dataset is that there are no reliable class labels which makes it hard to use for object recognition experiments. We created two sets of reliable labels. The CIFAR-10 set has 6000 examples of each of 10 classes and the CIFAR-100 set has 600 examples of each of 100 non-overlapping classes. Using these labels, we show that object recognition is signicantly improved by pre-training a layer of features on a large set of unlabeled tiny images.