Bayesian Robust Tensor Factorization for Incomplete Multiway Data

Bayesian Robust Tensor Factorization for Incomplete Multiway Data
复制标题

不完整多路数据的贝叶斯鲁棒张量分解

DOI:
10.1109/tnnls.2015.2423694
复制
发表时间:
2014-10
影响因子:
10.4
通讯作者:
Amari, Shun-Ichi
Amari, Shun-Ichi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhou, Guoxu;Zhang, Liqing;Cichocki, Andrzej;Amari, Shun-Ichi

文献摘要

参考文献

被引文献

相似文献

我们提出了一个同时存在缺失数据和离群点的稳健张量分解的生成模型。其目的是显式地推断出捕获全局信息的底层低CANDECOMP/PARAFAC(CP)秩张量和捕获局部信息(也被认为是离群值)的稀疏张量,从而提供关于丢失条目的稳健预测分布。低CP秩张量由多个潜在因素之间的多线性相互作用来建模,列稀疏性由分层先验来实施,而稀疏张量由学生t分布的分层视图来建模,该分层视图将单个超参数独立地与每个元素相关联。对于模型学习,我们在完全贝叶斯处理下发展了一种有效的变分推理,它可以有效地防止过拟合问题,并且随着数据量的增加而线性扩展。与已有的相关工作相比,我们的方法可以自动隐式地进行模型选择,而不需要调整参数。更具体地说,它可以发现CP排序的基本事实,并自动调整稀疏性诱导先验来适应各种类型的离群点。此外,在最大模型证据的意义下,可以优化低阶近似和稀疏表示之间的权衡。在合成数据集和真实数据集上的大量实验和与许多最先进的算法的比较从几个角度证明了我们的方法的优越性。
We propose a generative model for robust tensor factorization in the presence of both missing data and outliers. The objective is to explicitly infer the underlying low-CANDECOMP/PARAFAC (CP)-rank tensor capturing the global information and a sparse tensor capturing the local information (also considered as outliers), thus providing the robust predictive distribution over missing entries. The low-CP-rank tensor is modeled by multilinear interactions between multiple latent factors on which the column sparsity is enforced by a hierarchical prior, while the sparse tensor is modeled by a hierarchical view of Student-t distribution that associates an individual hyperparameter with each element independently. For model learning, we develop an efficient variational inference under a fully Bayesian treatment, which can effectively prevent the overfitting problem and scales linearly with data size. In contrast to existing related works, our method can perform model selection automatically and implicitly without the need of tuning parameters. More specifically, it can discover the groundtruth of CP rank and automatically adapt the sparsity inducing priors to various types of outliers. In addition, the tradeoff between the low-rank approximation and the sparse representation can be optimized in the sense of maximum model evidence. The extensive experiments and comparisons with many state-of-the-art algorithms on both synthetic and real-world data sets demonstrate the superiorities of our method from several perspectives.
DOI: --
发表时间: 2014-06
期刊: --
影响因子: --
作者:
Piyush Rai;Yingjian Wang;Shengbo Guo;Gary Chen;D. Dunson;L. Carin
通讯作者: Piyush Rai;Yingjian Wang;Shengbo Guo;Gary Chen;D. Dunson;L. Carin
通过多线性低n阶分解模型完成张量
DOI: 10.1016/j.neucom.2013.11.020
发表时间: 2014-06
期刊: Neurocomputing
影响因子: 6
作者:
Bin Cheng;Wuhong Wang;Yu-Jin Zhang;Bin Ran
通讯作者: Bin Ran
DOI: --
发表时间: 2011-08
期刊: arXiv: Learning
影响因子: --
作者:
Zenglin Xu;Feng Yan;Y. Qi
通讯作者: Zenglin Xu;Feng Yan;Y. Qi
DOI: 10.1007/s13042-011-0017-0
发表时间: 2011-03
影响因子: 5.6
作者:
Jie Li;Guan Han;Jing Wen;Xinbo Gao
通讯作者: Jie Li;Guan Han;Jing Wen;Xinbo Gao
DOI: 10.1137/120899066
发表时间: 2012-11
期刊: SIAM J. Matrix Anal. Appl.
影响因子: --
作者:
E. Allman;P. Jarvis;J. Rhodes;J. Sumner
通讯作者: E. Allman;P. Jarvis;J. Rhodes;J. Sumner