SINGULARITY, MISSPECIFICATION AND THE CONVERGENCE RATE OF EM

SINGULARITY, MISSPECIFICATION AND THE CONVERGENCE RATE OF EM
复制标题

DOI:
10.1214/19-aos1924
复制
发表时间:
2020-12-01
影响因子:
4.5
通讯作者:
Yu, Bin
Yu, Bin
中科院分区:
数学1区
文献类型:
--
作者:
Dwivedi, Raaz;Nhat Ho;Yu, Bin

文献摘要

被引文献

相似文献

最近的一系列工作分析了期望最大化(EM)算法在明确指定的设置下的行为,在该设置中,总体可能性围绕其最大化参数局部强凹。例子包括适当分离的高斯混合模型和线性回归的混合。我们考虑其中拟合分量的数量大于真实分布中的分量数量的过度指定的设置。这种错误的设置可能导致Fisher信息矩阵中的奇异性,进而导致基于nI.I.D.的最大似然估计器。D维的样本可以具有非标准的O((d/n)(1/4))收敛速度。针对d维高斯分布双组分混合模型的简单设置,我们研究了混合权值不同(不平衡情况)和相等(平衡情况)时EM算法的行为。我们的分析揭示了这两种情况的明显区别:在前一种情况下,EM算法以几何方法收敛到距离真实参数O((d/n)(1/2))的欧几里德距离处的一点,而在后一种情况下,收敛速度指数级地慢,并且不动点的精度要低得多。对这一特殊情况的分析需要引入一些新的技术:特别是,我们在相关的经验过程中利用了一种仔细的本地化形式,并开发了递归论证来逐步提高统计率。
A line of recent work has analyzed the behavior of the Expectation-Maximization (EM) algorithm in the well-specified setting, in which the population likelihood is locally strongly concave around its maximizing argument. Examples include suitably separated Gaussian mixture models and mixtures of linear regressions. We consider over-specified settings in which the number of fitted components is larger than the number of components in the true distribution. Such mis-specified settings can lead to singularity in the Fisher information matrix, and moreover, the maximum likelihood estimator based on n i.i.d. samples in d dimensions can have a nonstandard O((d/n)(1/4)) rate of convergence. Focusing on the simple setting of two-component mixtures fit to a d-dimensional Gaussian distribution, we study the behavior of the EM algorithm both when the mixture weights are different (unbalanced case), and are equal (balanced case). Our analysis reveals a sharp distinction between these two cases: in the former, the EM algorithm converges geomet- rically to a point at Euclidean distance of O((d/n)(1/2)) from the true parameter, whereas in the latter case, the convergence rate is exponentially slower, and the fixed point has a much lower O((d/n)(1/4)) accuracy. Analysis of this singular case requires the introduction of some novel techniques: in particular, we make use of a careful form of localization in the associated empirical process, and develop a recursive argument to progressively sharpen the statistical rate.