High-dimensional Motion Segmentation by Variational Autoencoder and Gaussian Processes

High-dimensional Motion Segmentation by Variational Autoencoder and Gaussian Processes
复制标题

通过变分自动编码器和高斯过程进行高维运动分割

DOI:
10.1109/iros40897.2019.8967987
复制
发表时间:
2019
期刊:
Proceedings of the International Conference on Intelligent Robots and Systems
影响因子:
--
通讯作者:
Takano Wataru
Takano Wataru
中科院分区:
--
文献类型:
--
作者:
Nagano Masatoshi;Nakamura Tomoaki;Nagai Takayuki;Mochihashi Daichi;Kobayashi Ichiro;Takano Wataru

文献摘要

相似文献

人类通过将连续的高维信息划分为诸如单词和运动单元等重要部分来感知连续的高维信息。我们认为,这种无监督分割对于机器人学习语言和运动等主题也很重要。为此,我们以前提出了一个层次的狄利克雷过程高斯过程隐半马尔可夫模型(HDP-GP-HSMM)。然而,该模型的一个重要缺点是它不能划分高维时间序列数据。此外,必须预先提取低维特征。分割在很大程度上依赖于特征的设计,设计有效的特征是困难的,特别是在高维数据的情况下。为了克服这个问题,本文提出了一种分层的Dirichlet过程-变分自编码器-高斯过程-隐半马尔可夫模型(HVGH)。建议HVGH的参数估计通过一个相互学习循环的变分自动编码器和我们以前提出的HDP-GP-HSMM。因此,HVGH可以从高维时间序列数据中提取特征,同时以无监督的方式将其划分为段。在一个实验中,我们使用了各种运动捕捉数据,以表明我们提出的模型估计正确的类的数量和更准确的段比基线方法。此外,我们表明,该方法可以学习潜在的空间适合分割。
Humans perceive continuous high-dimensional information by dividing it into significant segments such as words and units of motion. We believe that such unsupervised segmentation is also important for robots to learn topics such as language and motion. To this end, we previously proposed a hierarchical Dirichlet process-Gaussian process-hidden semi-Markov model (HDP-GP-HSMM). However, an important drawback to this model is that it cannot divide high-dimensional time-series data. Further, low-dimensional features must be extracted in advance. Segmentation largely depends on the design of features, and it is difficult to design effective features, especially in the case of high-dimensional data. To overcome this problem, this paper proposes a hierarchical Dirichlet process-variational autoencoder-Gaussian process-hidden semi-Markov model (HVGH). The parameters of the proposed HVGH are estimated through a mutual learning loop of the variational autoencoder and our previously proposed HDP-GP-HSMM. Hence, HVGH can extract features from high-dimensional time-series data, while simultaneously dividing it into segments in an unsupervised manner. In an experiment, we used various motion-capture data to show that our proposed model estimates the correct number of classes and more accurate segments than baseline methods. Moreover, we show that the proposed method can learn latent space suitable for segmentation.