Generalization Bounds in the Predict-then-Optimize Framework

Generalization Bounds in the Predict-then-Optimize Framework
复制标题

DOI:
10.1287/moor.2022.1330
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
Othman El Balghiti;Adam N. Elmachtoub;Paul Grigas;Ambuj Tewari
Othman El Balghiti;Adam N. Elmachtoub;Paul Grigas;Ambuj Tewari
中科院分区:
其他
文献类型:
--
作者:
Othman El Balghiti;Adam N. Elmachtoub;Paul Grigas;Ambuj Tewari

文献摘要

相似文献

预测-然后-优化框架在许多实际设置中是基本的:预测优化问题的未知参数,然后使用参数的预测值来解决问题。在这种情况下,一个自然损失函数是考虑由预测参数引起的决策成本,而不是参数的预测误差。该损失函数被称为智能预测然后优化(SPO)损失。在这项工作中,我们试图提供一个预测模型在训练数据上的性能在SPO损失的背景下对样本外的泛化程度的界限。因为SPO损失是非凸和非Lipschitz的,所以不适用于推导泛化界限的标准结果。我们首先基于Natarajan维导出界,在多面体可行域的情况下,极值点的个数至多是对数尺度,但在一般凸可行域的情况下,它与决策维线性相关。通过利用SPO损失函数的结构和可行域的一个关键属性,我们称之为强度属性,我们可以显著地改善对决策和特征维度的依赖。我们的方法和分析依赖于在不产生唯一最优解的有问题的预测周围放置保证金,然后在修正的保证金SPO损失函数是Lipschitz连续的情况下提供概括界。最后,我们刻画了强度的性质,并证明了对于强凸体和具有显式极点表示的多面体,修正的SPO损失都可以有效地计算出来。资金:O.El Balghiti感谢Rayens Capital的支持。埃尔马图布感谢美国国家科学基金会的支持[GRANT CMMI-1763000]。P.Grigas感谢美国国家科学基金会的支持[授予CCF-1755705和CMMI-1762744]。答:特瓦里感谢美国国家科学基金会[职业补助金IIS-1452099]和斯隆研究奖学金的支持。
The predict-then-optimize framework is fundamental in many practical settings: predict the unknown parameters of an optimization problem and then solve the problem using the predicted values of the parameters. A natural loss function in this environment is to consider the cost of the decisions induced by the predicted parameters in contrast to the prediction error of the parameters. This loss function is referred to as the smart predict-then-optimize (SPO) loss. In this work, we seek to provide bounds on how well the performance of a prediction model fit on training data generalizes out of sample in the context of the SPO loss. Because the SPO loss is nonconvex and non-Lipschitz, standard results for deriving generalization bounds do not apply. We first derive bounds based on the Natarajan dimension that, in the case of a polyhedral feasible region, scale at most logarithmically in the number of extreme points but, in the case of a general convex feasible region, have linear dependence on the decision dimension. By exploiting the structure of the SPO loss function and a key property of the feasible region, which we denote as the strength property, we can dramatically improve the dependence on the decision and feature dimensions. Our approach and analysis rely on placing a margin around problematic predictions that do not yield unique optimal solutions and then providing generalization bounds in the context of a modified margin SPO loss function that is Lipschitz continuous. Finally, we characterize the strength property and show that the modified SPO loss can be computed efficiently for both strongly convex bodies and polytopes with an explicit extreme point representation. Funding: O. El Balghiti thanks Rayens Capital for their support. A. N. Elmachtoub acknowledges the support of the National Science Foundation (NSF) [Grant CMMI-1763000]. P. Grigas acknowledges the support of NSF [Grants CCF-1755705 and CMMI-1762744]. A. Tewari acknowledges the support of the NSF [CAREER grant IIS-1452099] and a Sloan Research Fellowship.