Multiresolution Tensor Decompositions with Mode Hierarchies

Multiresolution Tensor Decompositions with Mode Hierarchies
复制标题

具有模式层次结构的多分辨率张量分解

DOI:
10.1145/2532169
复制
发表时间:
2014
期刊:
ACM Trans. Knowl. Discov. Data
影响因子:
--
通讯作者:
M. Sapino
M. Sapino
中科院分区:
--
文献类型:
--
作者:
C. Schifanella;K. Candan;M. Sapino

文献摘要

被引文献

相似文献

张量(多维数组)在从社交网络、传感器数据到互联网流量等应用中被广泛用于表示高维数据。多向数据分析技术,特别是张量分解,允许提取多向数据之间隐藏的相关性,因此是许多数据分析框架的关键组成部分。直观地说,这些算法可以被视为多向聚类方案,它在识别聚类、它们的权重以及每个数据元素的贡献时考虑数据的多个方面。不幸的是,用于拟合多向模型的算法通常是迭代的且非常耗时。在本文中,我们观察到,在许多应用中,存在关于一个或多个域维度的先验背景知识(或元数据)。这种元数据通常是以对给定数据面(或模式)的元素进行聚类的层次结构的形式存在。我们研究这种单模数据层次结构是否可用于提高张量分解过程的效率,而对最终分解质量没有重大影响。我们将每个域层次结构视为一种指南,以帮助按需提供张量中数据的更高或更低分辨率的视图,并且我们依靠这些由元数据诱导的多分辨率张量表示来开发一种张量分解的多分辨率方法。在本文中,我们专注于基于交替最小二乘法(ALS)实现两种最重要的分解模型,例如平行因子(PARAFAC,它将一个张量分解为一个对角张量和一组因子矩阵)和塔克(它产生一个核心张量和一组维度子空间矩阵作为结果)。实验结果表明,当可用的元数据被用作一个粗略的指南时,所提出的多分辨率方法有助于拟合PARAFAC和塔克模型,在执行时间和内存消耗方面有一致的(在不同参数设置下)节省,同时保持分解的质量。
Tensors (multidimensional arrays) are widely used for representing high-order dimensional data, in applications ranging from social networks, sensor data, and Internet traffic. Multiway data analysis techniques, in particular tensor decompositions, allow extraction of hidden correlations among multiway data and thus are key components of many data analysis frameworks. Intuitively, these algorithms can be thought of as multiway clustering schemes, which consider multiple facets of the data in identifying clusters, their weights, and contributions of each data element. Unfortunately, algorithms for fitting multiway models are, in general, iterative and very time consuming. In this article, we observe that, in many applications, there is a priori background knowledge (or metadata) about one or more domain dimensions. This metadata is often in the form of a hierarchy that clusters the elements of a given data facet (or mode). We investigate whether such single-mode data hierarchies can be used to boost the efficiency of tensor decomposition process, without significant impact on the final decomposition quality. We consider each domain hierarchy as a guide to help provide higher- or lower-resolution views of the data in the tensor on demand and we rely on these metadata-induced multiresolution tensor representations to develop a multiresolution approach to tensor decomposition. In this article, we focus on an alternating least squares (ALS)--based implementation of the two most important decomposition models such as the PARAllel FACtors (PARAFAC, which decomposes a tensor into a diagonal tensor and a set of factor matrices) and the Tucker (which produces as result a core tensor and a set of dimension-subspaces matrices). Experiment results show that, when the available metadata is used as a rough guide, the proposed multiresolution method helps fit both PARAFAC and Tucker models with consistent (under different parameters settings) savings in execution time and memory consumption, while preserving the quality of the decomposition.