A Non-Asymptotic Analysis of Stochastic Mirror Descent for Non-Convex Learning
A Non-Asymptotic Analysis of Stochastic Mirror Descent for Non-Convex Learning
批准号:
2444063
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
许多最先进的机器学习技术在很大程度上依赖于对某些非凸目标的优化。通常,这一过程的计算成本很高,并且需要使用使用人工噪声的随机近似或正则化技术。例如,深度学习中有各种技术,它们将人工噪声应用于数据、参数或更新方向,所有这些都已被证明鼓励更快的收敛和更好的泛化。虽然许多这些技术已经通过反复的实验和测试证明了它们的有效性,但验证和理解它们的理论框架仍处于起步阶段。另一种用于克服上述计算困难的技术是镜像下降框架,它通过利用所考虑问题的已知几何性质来实现这一点。它在凸优化中非常流行,因为在这种情况下,该算法有许多可证明的优点。其对应的随机镜像下降在非凸学习中迅速流行起来。但同样,在这种情况下的理论框架是不发达的。通过我们的项目,我们试图更好地理解这些技术在大规模学习环境中的统计和计算特性。具体地说,我们感兴趣的是定量比较这些技术的一致性和泛化性能,并看看它们如何随着问题的变得更具挑战性而变化。指导我们的大部分方法将专注于非渐近分析,即与将量发挥到无限极限的分析进行比较。我们相信,这对于大规模学习中的应用是至关重要的,在大规模学习中,只有有限的迭代次数可以发生,并且参数数量对性能的影响是关键的。有了这一点,我们就可以使用一套广泛的强大技术来分析扩散。近年来,这种方法越来越受欢迎,并在学习环境中产生了有趣的结果。该项目还将大量借鉴随机微分方程和高维统计的研究。该项目属于EPSRC统计和应用概率领域。
英文摘要
Many state of the art machine learning techniques depend heavily on the optimisation of some non-convex objective. Often, this procedure is computationally expensive and requires techniques that employ stochastic approximations or regularisation using artificial noise. For instance, there are a variety of techniques in deep learning that use artificial noise applied to the data, parameters or update direction, all of which have been shown to encourage faster convergence and improved generalisation. Though many of these techniques have demonstrated their validity through repeated experimentation and testing, the theoretical framework to validate and understand them is still in its infancy.Another technique which is used to overcome the aforementioned computational difficulties is the mirror descent framework, which does so by taking advantage of the known geometric properties specific to the problem in consideration. It is highly popular in convex optimisation as in this setting there are many provable advantages to this algorithm. Its stochastic counterpart, stochastic mirror descent, is rapidly gaining popularity in non-convex learning. But again, the theoretical framework in this setting is under-developed.Through our project, we seek to develop a better understanding of the statistical and computational properties of these techniques in the large-scale learning setting. Specifically, we are interested in quantitively comparing the consistency and generalisation performance of these techniques and seeing how they change as the problem grows more challenging.Guiding much of our methodology will be the focus on non-asymptotic analysis, that is compared to analyses in which quantities are taken to their infinite limits. We believe this is essential for applications in large-scale learning where only a limited number of iterations can take place and the effect of the number of parameters on performance is of key interest.Fundamentally, our approach will be based on approximating the iterative optimisation procedure with stochastic processes that evolve continuously in time, specifically diffusion processes. With this, we gain access to the broad suite of powerful techniques for analysing diffusions. This approach has seen a growing popularity in recent years and has already yielded interesting results in the learning setting. The project will also draw heavily from the study of stochastic differential equations as well as high dimensional statistics.This project falls within the EPSRC Statistics and Applied probability area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金