The Gaussian Process Prior VAE for Interpretable Latent Dynamics from Pixels

The Gaussian Process Prior VAE for Interpretable Latent Dynamics from Pixels
复制标题

用于解释像素潜在动态的高斯过程先验 VAE

DOI:
--
复制
发表时间:
2019
期刊:
Symposium on Advances in Approximate Bayesian Inference
影响因子:
--
通讯作者:
Michael Pearce
Michael Pearce
中科院分区:
--
文献类型:
--
作者:
Michael Pearce

文献摘要

被引文献

相似文献

我们考虑了包含运动对象的视频的低维、可解释、潜在状态的无监督学习问题。从像素中提取可解释动态的问题已经通过图形/状态空间模型(Fraccaro等人,2017;Lin等人,2018;Pearce等人,2018;Chiappa和Paquet,2019)的视角得到了广泛的考虑,这些模型利用马尔可夫结构进行廉价计算,并构造先验来强制对潜在表征进行可解释。我们通过摒弃马尔可夫结构,朝着扩展这些方法迈出了一步;受到高斯过程动力学模型(Wang等人,2006年)的启发,我们取而代之的是最近提出的高斯过程优先变分自动编码器(Casale等人,2018年),用于学习可解释的潜在动力学。我们描述了该模型,并在一个合成数据集上进行了实验,发现该模型可靠地重建了呈现U形转弯和循环的平滑动力学。我们还观察到,与先前的工作相反,该模型可以在没有任何β退火或训练参数冻融的情况下进行训练,尽管用于略有不同的用例,其中经常需要特定于应用的训练技巧。
We consider the problem of unsupervised learning of a low dimensional, interpretable, latent state of a video containing a moving object. The problem of distilling interpretable dynamics from pixels has been extensively considered through the lens of graphical/state space models (Fraccaro et al., 2017; Lin et al., 2018; Pearce et al., 2018; Chiappa and Paquet, 2019) that exploit Markov structure for cheap computation and structured priors for enforcing interpretability on latent representations. We take a step towards extending these approaches by discarding the Markov structure; inspired by Gaussian process dynamical models (Wang et al., 2006), we instead repurpose the recently proposed Gaussian Process Prior Variational Autoencoder (Casale et al., 2018) for learning interpretable latent dynamics. We describe the model and perform experiments on a synthetic dataset and see that the model reliably reconstructs smooth dynamics exhibiting U-turns and loops. We also observe that this model may be trained without any β annealing or freeze-thaw of training parameters in contrast to previous works, albeit for slightly different use cases, where application specific training tricks are often required.