End-to-End Stochastic Optimization with Energy-Based Model

End-to-End Stochastic Optimization with Energy-Based Model
复制标题

DOI:
10.48550/arxiv.2211.13837
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Lingkai Kong;Jiaming Cui;Yuchen Zhuang;Rui Feng;B. Prakash;Chao Zhang
Lingkai Kong;Jiaming Cui;Yuchen Zhuang;Rui Feng;B. Prakash;Chao Zhang
中科院分区:
其他
文献类型:
--
作者:
Lingkai Kong;Jiaming Cui;Yuchen Zhuang;Rui Feng;B. Prakash;Chao Zhang

文献摘要

相似文献

决策中心学习(Decision-focused learning, DFL)是近年来提出的一种用于求解未知参数随机优化问题的方法。通过将预测建模与隐式微优化层相结合,DFL显示出优于标准两阶段预测-优化管道的性能。然而,大多数现有的DFL方法仅适用于凸问题或易于松弛为凸问题的非凸问题子集。此外,由于每次训练迭代都需要通过优化问题进行求解和微分,因此训练效率低下。本文提出了一种基于能量模型的通用高效DFL随机优化方法SO-EBM。SO-EBM不是依靠KKT条件来推导隐式优化层,而是使用基于能量函数的可微优化层显式地参数化原始优化问题。为了更好地近似优化景观,我们提出了一个耦合训练目标,该目标使用最大似然损失来捕获最佳位置,并使用基于分布的正则化器来捕获整体能源景观。最后,我们提出了一种基于高斯混合建议的自归一化重要性采样器的SO-EBM有效训练方法。在电力调度、COVID-19资源分配和非凸对抗性安全博弈三种应用中对SO-EBM进行了评估,验证了SO-EBM的有效性和高效性。
Decision-focused learning (DFL) was recently proposed for stochastic optimization problems that involve unknown parameters. By integrating predictive modeling with an implicitly differentiable optimization layer, DFL has shown superior performance to the standard two-stage predict-then-optimize pipeline. However, most existing DFL methods are only applicable to convex problems or a subset of nonconvex problems that can be easily relaxed to convex ones. Further, they can be inefficient in training due to the requirement of solving and differentiating through the optimization problem in every training iteration. We propose SO-EBM, a general and efficient DFL method for stochastic optimization using energy-based models. Instead of relying on KKT conditions to induce an implicit optimization layer, SO-EBM explicitly parameterizes the original optimization problem using a differentiable optimization layer based on energy functions. To better approximate the optimization landscape, we propose a coupled training objective that uses a maximum likelihood loss to capture the optimum location and a distribution-based regularizer to capture the overall energy landscape. Finally, we propose an efficient training procedure for SO-EBM with a self-normalized importance sampler based on a Gaussian mixture proposal. We evaluate SO-EBM in three applications: power scheduling, COVID-19 resource allocation, and non-convex adversarial security game, demonstrating the effectiveness and efficiency of SO-EBM.