Collective Tensor Completion with Multiple Heterogeneous Side Information

Collective Tensor Completion with Multiple Heterogeneous Side Information
复制标题

DOI:
10.1109/bigdata47090.2019.9006072
复制
发表时间:
2019-12
期刊:
2019 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Huiyuan Chen;Jing Li
Huiyuan Chen;Jing Li
中科院分区:
其他
文献类型:
--
作者:
Huiyuan Chen;Jing Li

文献摘要

相似文献

张量补全已经成功地应用到许多实际应用中。在各种各样的情况下,许多学习任务中使用的数据都是高维的,通常是从多个异构源中提取的。因此,数据可以由一个主张量和由多视图侧信息或元数据生成的多个矩阵来表示。张量和矩阵的联合分析对于更好地理解这些多重异构源之间的潜在关系具有很大的潜力。现有的张量补全方法是利用单一视图侧信息恢复部分已知张量的缺失元素,可以产生可解释的大规模数据集结果。然而,目前这些方法的局限性在于缺乏对多视图异构数据的建模和对张量低秩特性的适当学习。在本研究中,我们通过开发一种新颖的集体张量补全方法来填补这一空白,该方法紧密融合了多视图异构数据源。我们的方法通过张量-矩阵耦合分解,从主张量和多侧矩阵中挖掘出特殊的共同潜在结构,其中共同潜在结构可以紧凑地表示所有数据。此外,由于张量的离散性,其秩估计是一项具有挑战性的任务。代替常用的迹范数或核范数来逼近秩,我们直接在潜在结构上使用Schatten p-范数来更好地逼近秩并增强其对噪声的鲁棒性。
Tensor completion has been successfully applied to many real-world applications. In a wide variety of situations, data utilized in many learning tasks are of high dimensions, usually extracted from multiple heterogeneous sources. Therefore, data can be represented by a primary tensor and multiple matrices generated from multi-view side information or metadata. Joint analysis of tensors and matrices has great potential to gain better understanding of the underlying relationships among these multiple heterogeneous sources. The existing tensor completion methods, which recover the missing elements of a partially known tensor with single view side information, can yield interpretable results for large-scale datasets. However, their limitations up to now are lack of modeling multi-view heterogeneous data and suitably learning the low-rank property of tensor. In this study, we fill this gap by developing a novel collective tensor completion method, which tightly fuses multi-view heterogeneous data sources. Our method exploits special common latent structures from the primary tensor and multiple side matrices through coupled tensor-matrix decomposition, in which the common latent structures can compactly represent all the data. In addition, rank estimation of a tensor is a challenging task due to its discrete nature. Instead of approximating the rank by widely used trace norm or nuclear norm, we directly utilize Schatten p-norm on the latent structures to better approximate the rank and to enhance its robustness to noise.