A Stochastic Subgradient Method for Distributionally Robust Non-convex and Non-smooth Learning

A Stochastic Subgradient Method for Distributionally Robust Non-convex and Non-smooth Learning
复制标题

DOI:
10.1007/s10957-022-02063-6
复制
发表时间:
2022-07
影响因子:
1.9
通讯作者:
M. Gürbüzbalaban;A. Ruszczynski;Landi Zhu
M. Gürbüzbalaban;A. Ruszczynski;Landi Zhu
中科院分区:
数学3区
文献类型:
--
作者:
M. Gürbüzbalaban;A. Ruszczynski;Landi Zhu

文献摘要

相似文献

我们认为,在统计学习中产生的随机优化问题,鲁棒性是相对于模糊性的基础数据分布的分布鲁棒性制定。我们的配方建立在风险厌恶的优化技术和一致的风险措施的理论。它使用平均半确定性风险来量化不确定性,使我们能够计算出对人口数据分布扰动具有鲁棒性的解决方案。我们考虑了广义可微损失函数,可以是非凸的和非光滑的,涉及向上和向下的尖点,我们开发了一个有效的随机次梯度方法的分布鲁棒性问题,这样的功能。我们证明了它收敛到一个点满足最优性条件。据我们所知,这是第一个严格的收敛保证的广义可微非凸和非光滑分布鲁棒随机优化的背景下的方法。我们的方法允许控制所需的鲁棒性水平,很少额外的计算成本相比,人口风险最小化与随机梯度方法。我们还说明了我们的算法在凸和非凸监督学习问题中产生的真实的数据集上的性能。
We consider a distributionally robust formulation of stochastic optimization problems arising in statistical learning, where robustness is with respect to ambiguity in the underlying data distribution. Our formulation builds on risk-averse optimization techniques and the theory of coherent risk measures. It uses mean–semideviation risk for quantifying uncertainty, allowing us to compute solutions that are robust against perturbations in the population data distribution. We consider a broad class of generalized differentiable loss functions that can be non-convex and non-smooth, involving upward and downward cusps, and we develop an efficient stochastic subgradient method for distributionally robust problems with such functions. We prove that it converges to a point satisfying the optimality conditions. To our knowledge, this is the first method with rigorous convergence guarantees in the context of generalized differentiable non-convex and non-smooth distributionally robust stochastic optimization. Our method allows for the control of the desired level of robustness with little extra computational cost compared to population risk minimization with stochastic gradient methods. We also illustrate the performance of our algorithm on real datasets arising in convex and non-convex supervised learning problems.