First-order and Stochastic Optimization Methods for Machine Learning

First-order and Stochastic Optimization Methods for Machine Learning
复制标题

DOI:
10.1007/978-3-030-39568-1
复制
发表时间:
2020
期刊:
JAIDS Journal of Acquired Immune Deficiency Syndromes
影响因子:
--
通讯作者:
Guanghui Lan
Guanghui Lan
中科院分区:
其他
文献类型:
--
作者:
Guanghui Lan

文献摘要

被引文献

相似文献

从一开始,优化就在数据科学中扮演着至关重要的角色。许多统计和机器学习模型的分析和求解方法依赖于优化。最近对计算数据分析优化的兴趣的激增也伴随着一些重大的挑战。高问题维度、大数据量、固有的不确定性、不可避免的非凸性,以及对实时、有时在分布式环境下解决这些问题的日益增长的需求,这些都为现有的优化方法创造了相当大的障碍。在过去的10年左右的时间里,优化算法的设计和分析已经取得了重大进展,以应对其中的一些挑战。然而,它们散布在一些不同学科的大量文献中。这些进展缺乏系统的处理,使得年轻的研究人员越来越难以踏入这一领域,建立必要的基础,了解当前的技术水平,并推动这一令人兴奋的研究领域的前沿。在这本书中,我试图将这些最新进展中的一些放在稍微更有条理的方式中。我主要关注已经被广泛应用或(在我看来)可能具有应用于大规模机器学习和数据分析的应用潜力的优化算法。这些方法包括相当多的一阶方法、随机优化方法、随机化方法和分布方法、非凸随机优化方法、无投影方法以及算子滑动和分散方法。我的目标是介绍能够在不同设置下提供最佳性能保证的基本算法方案。在讨论这些算法之前,我简要介绍了几个流行的机器学习模型,以启发读者,并回顾了一些重要的优化理论,为读者,特别是初学者,提供了良好的理论基础。这本书的目标读者包括对优化方法及其在机器学习或机器智能中的应用感兴趣的研究生和大四本科生。它也可以作为更资深的研究人员的参考书。这本书的初稿已经被用作佐治亚学院的高级本科班和博士班的课本
Since its beginning, optimization has played a vital role in data science. The analysis and solution methods for many statistical and machine learning models rely on optimization. The recent surge of interest in optimization for computational data analysis also comes with a few significant challenges. The high problem dimensionality, large data volumes, inherent uncertainty, unavoidable nonconvexity, together with the increasing need to solve these problems in real time and sometimes under a distributed setting all contribute to create a considerable set of obstacles for existing optimization methodologies.During the past 10 years or so, significant progresses have been made in the design and analysis of optimization algorithms to tackle some of these challenges. Nevertheless, they were scattered in a large body of literature across a few different disciplines. The lack of a systematic treatment for these progresses makes it more and more difficult for young researchers to step into this field, build up the necessary foundation, understand the current state of the art, and push forward the frontier of this exciting research area. In this book, I attempt to put some of these recent progresses into a slightly more organized manner. I mainly focus on the optimization algorithms that have been widely applied or may have the applied potential (from my perspective) to large-scale machine learning and data analysis. These include quite a few first-order methods, stochastic optimization methods, randomized and distributed methods, nonconvex stochastic optimization methods, projection-free methods, and operator sliding and decentralized methods. My goal is to introduce the basic algorithmic schemes that can provide the best performance guarantees under different settings. Before discussing these algorithms, I do provide a brief introduction to a few popular machine learning models to inspire the readers and also review some important optimization theory to equip the readers, especially the beginners, with a good theoretic foundation. The target audience of this book includes the graduate students and senior undergraduate students who are interested in optimization methods and their applications in machine learning or machine intelligence. It can also be used as a reference book for more senior researchers. The initial draft of this book has been used as the text for a senior undergraduate class and a Ph. D. class here at Georgia Institute