The scaling limit of high-dimensional online independent component analysis

The scaling limit of high-dimensional online independent component analysis
复制标题

DOI:
10.1088/1742-5468/ab39d6
复制
发表时间:
2017-10
期刊:
Journal of Statistical Mechanics: Theory and Experiment
影响因子:
--
通讯作者:
Chuang Wang;Yue M. Lu
Chuang Wang;Yue M. Lu
中科院分区:
其他
文献类型:
--
作者:
Chuang Wang;Yue M. Lu

文献摘要

相似文献

我们分析了高维缩放限制下独立分量分析在线算法的动态。由于环境维度趋于无穷大,并且在适当的时间缩放下,我们表明目标特征向量的时变联合经验测量和算法提供的估计将弱地收敛到确定性测量值过程,该过程可以表征为非线性偏微分方程的唯一解。该偏微分方程涉及两个空间变量和一个时间变量,可以有效地获得数值解。这些解决方案提供了有关 ICA 算法性能的详细信息,因为许多实际性能指标都是联合经验测量的函数。数值模拟表明,即使对于中等尺寸,我们的渐近分析也是准确的。除了提供了解算法性能的工具外,我们的偏微分方程分析还提供了有用的见解。特别是,在高维限制下,与算法相关的原始耦合动力学将渐近“解耦”,每个坐标通过随机梯度下降独立求解一维有效最小化问题。利用这种洞察力来设计新算法,以实现计算效率和统计效率之间的最佳权衡,可能会成为未来研究的一个有趣方向。
We analyze the dynamics of an online algorithm for independent component analysis in the high-dimensional scaling limit. As the ambient dimension tends to infinity, and with proper time scaling, we show that the time-varying joint empirical measure of the target feature vector and the estimates provided by the algorithm will converge weakly to a deterministic measured-valued process that can be characterized as the unique solution of a nonlinear PDE. Numerical solutions of this PDE, which involves two spatial variables and one time variable, can be efficiently obtained. These solutions provide detailed information about the performance of the ICA algorithm, as many practical performance metrics are functionals of the joint empirical measures. Numerical simulations show that our asymptotic analysis is accurate even for moderate dimensions. In addition to providing a tool for understanding the performance of the algorithm, our PDE analysis also provides useful insight. In particular, in the high-dimensional limit, the original coupled dynamics associated with the algorithm will be asymptotically ‘decoupled’, with each coordinate independently solving a 1D effective minimization problem via stochastic gradient descent. Exploiting this insight to design new algorithms for achieving optimal trade-offs between computational and statistical efficiency may prove an interesting line of future research.