Randomized Algorithms for Low-Rank Tensor Decompositions in the Tucker Format

Randomized Algorithms for Low-Rank Tensor Decompositions in the Tucker Format
复制标题

DOI:
10.1137/19m1261043
复制
发表时间:
2020-01-01
影响因子:
3.6
通讯作者:
Kilmer, Misha E.
Kilmer, Misha E.
中科院分区:
数学2区
文献类型:
--
作者:
Minster, Rachel;Saibaba, Arvind K.;Kilmer, Misha E.

文献摘要

被引文献

相似文献

数据科学和科学计算中的许多应用涉及存储和操作成本高昂的大规模数据集。然而,这些数据集具有固有的多维结构,可以用来压缩和存储数据集在一个适当的张量格式。近年来,随机矩阵方法已被用来有效和准确地计算低秩矩阵分解。受此成功的启发,我们开发了随机算法的张量分解的塔克表示。具体而言,我们提出了两个著名的压缩算法,即HOSVD和STHOSVD,并在使用这两种算法的错误的详细概率分析的随机版本。我们还开发了这些算法的变体,以解决大规模数据集带来的特定挑战。第一种变体自适应地找到满足给定容限的低秩表示,并且当目标秩事先不知道时是有益的。第二种变体保留了原始张量的结构,并且对于难以加载到内存中的大型稀疏张量非常有益。我们考虑了几个不同的数据集,我们的数值实验:合成测试张量和现实的应用程序,如Olivetti数据库中的面部图像样本的压缩和安然电子邮件数据集中的字数。
Many applications in data science and scientific computing involve large-scale datasets that are expensive to store and manipulate. However, these datasets possess inherent multidimensional structure that can be exploited to compress and store the dataset in an appropriate tensor format. In recent years, randomized matrix methods have been used to efficiently and accurately compute low-rank matrix decompositions. Motivated by this success, we develop randomized algorithms for tensor decompositions in the Tucker representation. Specifically, we present randomized versions of two well-known compression algorithms, namely, HOSVD and STHOSVD, and a detailed probabilistic analysis of the error in using both algorithms. We also develop variants of these algorithms that tackle specific challenges posed by large-scale datasets. The first variant adaptively finds a low-rank representation satisfying a given tolerance, and it is beneficial when the target rank is not known in advance. The second variant preserves the structure of the original tensor and is beneficial for large sparse tensors that are difficult to load in memory. We consider several different datasets for our numerical experiments: synthetic test tensors and realistic applications such as the compression of facial image samples in the Olivetti database and word counts in the Enron email dataset.