Computation of the maximum likelihood estimator in low-rank factor analysis

Computation of the maximum likelihood estimator in low-rank factor analysis
复制标题

DOI:
10.1007/s10107-019-01370-7
复制
发表时间:
2018-01
影响因子:
2.7
通讯作者:
K. Khamaru;R. Mazumder
K. Khamaru;R. Mazumder
中科院分区:
数学2区
文献类型:
--
作者:
K. Khamaru;R. Mazumder

文献摘要

相似文献

因子分析是一种经典的多元降维技术,广泛应用于统计学、计量经济学和数据科学。因子分析的估计通常通过最大似然原理进行,该原理寻求在正定协方差矩阵可以分解为低秩正半定矩阵和具有非负项的对角矩阵之和的假设下最大化高斯似然。这导致了一个具有挑战性的秩约束非凸优化问题,对此可用的可靠计算算法非常少。我们将低秩最大似然因子分析任务重新表述为非线性非光滑半定优化问题,研究该重新表述的各种结构特性;并提出基于凸优化差异的快速且可扩展的算法。我们的方法具有计算保证,可以优雅地扩展到大型问题,适用于样本协方差矩阵秩不足的情况,并适应最大似然问题的变体,并对模型参数进行附加约束。我们的数值实验验证了我们的方法相对于现有最先进的最大似然因子分析方法的有用性。
Factor analysis is a classical multivariate dimensionality reduction technique popularly used in statistics, econometrics and data science. Estimation for factor analysis is often carried out via the maximum likelihood principle, which seeks to maximize the Gaussian likelihood under the assumption that the positive definite covariance matrix can be decomposed as the sum of a low-rank positive semidefinite matrix and a diagonal matrix with nonnegative entries. This leads to a challenging rank constrained nonconvex optimization problem, for which very few reliable computational algorithms are available. We reformulate the low-rank maximum likelihood factor analysis task as a nonlinear nonsmooth semidefinite optimization problem, study various structural properties of this reformulation; and propose fast and scalable algorithms based on difference of convex optimization. Our approach has computational guarantees, gracefully scales to large problems, is applicable to situations where the sample covariance matrix is rank deficient and adapts to variants of the maximum likelihood problem with additional constraints on the model parameters. Our numerical experiments validate the usefulness of our approach over existing state-of-the-art approaches for maximum likelihood factor analysis.