Infinite Tucker Decomposition: Nonparametric Bayesian Models for Multiway Data Analysis

Infinite Tucker Decomposition: Nonparametric Bayesian Models for Multiway Data Analysis
复制标题

DOI:
--
复制
发表时间:
2011-08
期刊:
arXiv: Learning
影响因子:
--
通讯作者:
Zenglin Xu;Feng Yan;Y. Qi
Zenglin Xu;Feng Yan;Y. Qi
中科院分区:
其他
文献类型:
--
作者:
Zenglin Xu;Feng Yan;Y. Qi

文献摘要

被引文献

相似文献

张量分解是一种强大的多向数据分析计算工具。许多流行的张量分解方法-如Tucker分解和CANDECOMP/PARAFAC(CP)-相当于多线性因式分解。它们不足以模拟(I)数据实体之间的复杂交互,(Ii)各种数据类型(例如,丢失数据和二进制数据),以及(Iii)有噪声的观测和离群值。为了解决这些问题,我们提出了张量-变量潜在非参数贝叶斯模型,并结合有效的推理方法,用于多路数据分析。我们将这些模型命名为InfTucker。利用这些InfTucker,我们在无限的特征空间中进行Tucker分解。与经典张量分解模型不同,我们的新方法在概率框架中处理连续和二进制数据。与以往关于矩阵和张量的贝叶斯模型不同,我们的模型是基于具有非线性协方差函数的潜在高斯过程或$t$过程。为了有效地从数据中学习InfTucker,我们开发了一种关于张量的变分推理技术。与经典实现相比,新技术将时间和空间复杂度降低了几个数量级。我们在化学计量学和社会网络数据集上的实验结果表明,我们的新模型获得了比最先进的张量分解更高的预测精度
Tensor decomposition is a powerful computational tool for multiway data analysis. Many popular tensor decomposition approaches---such as the Tucker decomposition and CANDECOMP/PARAFAC (CP)---amount to multi-linear factorization. They are insufficient to model (i) complex interactions between data entities, (ii) various data types (e.g. missing data and binary data), and (iii) noisy observations and outliers. To address these issues, we propose tensor-variate latent nonparametric Bayesian models, coupled with efficient inference methods, for multiway data analysis. We name these models InfTucker. Using these InfTucker, we conduct Tucker decomposition in an infinite feature space. Unlike classical tensor decomposition models, our new approaches handle both continuous and binary data in a probabilistic framework. Unlike previous Bayesian models on matrices and tensors, our models are based on latent Gaussian or $t$ processes with nonlinear covariance functions. To efficiently learn the InfTucker from data, we develop a variational inference technique on tensors. Compared with classical implementation, the new technique reduces both time and space complexities by several orders of magnitude. Our experimental results on chemometrics and social network datasets demonstrate that our new models achieved significantly higher prediction accuracy than the most state-of-art tensor decomposition