Composite Difference-Max Programs for Modern Statistical Estimation Problems

Composite Difference-Max Programs for Modern Statistical Estimation Problems
复制标题

DOI:
10.1137/18m117337x
复制
发表时间:
2018-03
期刊:
SIAM J. Optim.
影响因子:
--
通讯作者:
Ying Cui;J. Pang;B. Sen
Ying Cui;J. Pang;B. Sen
中科院分区:
其他
文献类型:
--
作者:
Ying Cui;J. Pang;B. Sen

文献摘要

被引文献

相似文献

许多现代统计估计问题由三个主要组成部分定义:假设输出变量对输入特征的依赖性的统计模型;测量观测输出和模型预测输出之间误差的损失函数;以及控制模型中过拟合和/或变量选择的正则化器。我们研究了这个通用的统计估计问题的抽样版本,其中模型参数估计的经验风险最小化,这涉及到最小化的经验平均损失函数的数据点加权模型正则化。在我们的设置中,我们允许上面讨论的所有三个分量函数都是凸差分(dc)类型,并使用大量常用的例子来说明它们,包括连续分段仿射回归和深度学习中的例子(其中激活函数是分段仿射的)。我们描述了一个非单调优化最小化(MM)算法解决统一的非凸,不可微的优化问题,这是制定为一个特殊结构的复合直流程序的逐点最大值型,并提出收敛结果的方向固定的解决方案。提出了一种有效的半光滑牛顿法来求解MM子问题的对偶。数值结果表明,所提出的算法的有效性和连续分段仿射回归优于标准线性模型。
Many modern statistical estimation problems are defined by three major components: a statistical model that postulates the dependence of an output variable on the input features; a loss function measuring the error between the observed output and the model predicted output; and a regularizer that controls the overfitting and/or variable selection in the model. We study the sampling version of this generic statistical estimation problem where the model parameters are estimated by empirical risk minimization, which involves the minimization of the empirical average of the loss function at the data points weighted by the model regularizer. In our setup we allow all three component functions discussed above to be of the difference-of-convex (dc) type and illustrate them with a host of commonly used examples, including those in continuous piecewise affine regression and in deep learning (where the activation functions are piecewise affine). We describe a nonmonotone majorization-minimization (MM) algorithm for solving the unified nonconvex, nondifferentiable optimization problem which is formulated as a specially structured composite dc program of the pointwise max type, and present convergence results to a directional stationary solution. An efficient semismooth Newton method is proposed to solve the dual of the MM subproblems. Numerical results are presented to demonstrate the effectiveness of the proposed algorithm and the superiority of continuous piecewise affine regression over the standard linear model.