Understanding Deflation Process in Over-parametrized Tensor Decomposition

Understanding Deflation Process in Over-parametrized Tensor Decomposition
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Rong Ge;Y. Ren;Xiang Wang;Mo Zhou
Rong Ge;Y. Ren;Xiang Wang;Mo Zhou
中科院分区:
其他
文献类型:
--
作者:
Rong Ge;Y. Ren;Xiang Wang;Mo Zhou

文献摘要

被引文献

相似文献

本文研究了梯度流的超参数化张量分解训练动力学问题。从经验上看,这种训练过程通常是先拟合较大的分量,然后发现较小的分量,这类似于张量分解算法中常用的张量紧缩过程。我们证明了对于正交可分解张量,稍加修改的梯度流将遵循张量收缩过程并恢复所有张量分量。我们的证明表明,对于正交张量,梯度流动力学与矩阵设置中的贪婪低秩学习相似,这是理解低秩张量的过参数化模型的隐式正则化效应的第一步。
In this paper we study the training dynamics for gradient flow on over-parametrized tensor decomposition problems. Empirically, such training process often first fits larger components and then discovers smaller components, which is similar to a tensor deflation process that is commonly used in tensor decomposition algorithms. We prove that for orthogonally decomposable tensor, a slightly modified version of gradient flow would follow a tensor deflation process and recover all the tensor components. Our proof suggests that for orthogonal tensors, gradient flow dynamics works similarly as greedy low-rank learning in the matrix setting, which is a first step towards understanding the implicit regularization effect of over-parametrized models for low-rank tensors.