Learning to Represent Spatial Transformations with Factored Higher-Order Boltzmann Machines

Learning to Represent Spatial Transformations with Factored Higher-Order Boltzmann Machines
复制标题

DOI:
10.1162/neco.2010.01-09-953
复制
发表时间:
2010-06-01
期刊:
影响因子:
2.9
通讯作者:
Hinton, Geoffrey E.
Hinton, Geoffrey E.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Memisevic, Roland;Hinton, Geoffrey E.

文献摘要

被引文献

相似文献

为了允许受限玻尔兹曼机的隐藏单元对两个连续图像之间的变换进行建模,Memisevic和欣顿(2007)引入了三向乘法交互,该交互使用第一图像中像素的强度作为第二图像中像素和隐藏单元之间的学习对称权重的乘法增益。这就产生了许多立方的参数,它们形成了一个三维的相互作用张量。我们描述了这个相互作用张量的低秩近似,它使用因子之和,每个因子都是三向外积。这种近似允许有效学习较大图像块之间的变换。由于每个因子都可以被视为图像滤波器,因此模型作为一个整体学习最佳滤波器对,以有效地表示变换。我们演示了从各种合成和真实的图像序列的最佳滤波器对的学习。我们还展示了如何学习图像变换,使模型能够执行简单的视觉类比任务,我们展示了一个完全无监督的网络如何以与人类相同的方式感知透明点图案的多种运动。
To allow the hidden units of a restricted Boltzmann machine to model the transformation between two successive images, Memisevic and Hinton (2007) introduced three-way multiplicative interactions that use the intensity of a pixel in the first image as a multiplicative gain on a learned, symmetric weight between a pixel in the second image and a hidden unit. This creates cubically many parameters, which form a three-dimensional interaction tensor. We describe a low-rank approximation to this interaction tensor that uses a sum of factors, each of which is a three-way outer product. This approximation allows efficient learning of transformations between larger image patches. Since each factor can be viewed as an image filter, the model as a whole learns optimal filter pairs for efficiently representing transformations. We demonstrate the learning of optimal filter pairs from various synthetic and real image sequences. We also show how learning about image transformations allows the model to perform a simple visual analogy task, and we show how a completely unsupervised network trained on transformations perceives multiple motions of transparent dot patterns in the same way as humans.