Algebraic optimization of sequential decision problems

Algebraic optimization of sequential decision problems
复制标题

顺序决策问题的代数优化

DOI:
10.1016/j.jsc.2023.102241
复制
发表时间:
2024
影响因子:
0.7
通讯作者:
Rose, Kemal
Rose, Kemal
中科院分区:
数学2区
文献类型:
--
作者:
Dressler, Mareike;Garrote-López, Marina;Montúfar, Guido;Müller, Johannes;Rose, Kemal

文献摘要

参考文献

相似文献

研究了平稳随机策略集上有限部分可观察马尔可夫决策过程中期望长期回报的优化问题。在确定性观测(也称为状态聚合)的情况下,该问题相当于在二次约束下优化线性目标。我们将这个问题的可行集刻画为一阶矩阵仿射变体与多面体之积的交集。在此基础上,我们得到了优化问题的临界点个数的界。最后,我们进行了实验,在可行集的不同边界分量上求解KKT方程或Lagrange方程,并将结果与理论边界和其他约束优化方法进行了比较。
We study the optimization of the expected long-term reward in finite partially observable Markov decision processes over the set of stationary stochastic policies. In the case of deterministic observations, also known as state aggregation, the problem is equivalent to optimizing a linear objective subject to quadratic constraints. We characterize the feasible set of this problem as the intersection of a product of affine varieties of rank one matrices and a polytope. Based on this description, we obtain bounds on the number of critical points of the optimization problem. Finally, we conduct experiments in which we solve the KKT equations or the Lagrange equations over different boundary components of the feasible set, and compare the result to the theoretical bounds and to other constrained optimization methods.
随机博弈中的实代数工具
DOI: 10.1007/978-94-010-0189-2_6
发表时间: 2003
期刊: ArXiv
影响因子: --
作者:
A. Neyman
通讯作者: A. Neyman
使用二次约束线性程序求解 POMDP
DOI: --
发表时间: 2006
期刊: Adaptive Agents and Multi-Agent Systems
影响因子: --
作者:
Chris Amato;D. Bernstein;S. Zilberstein
通讯作者: S. Zilberstein
马尔可夫决策过程的进一步实际应用
DOI: 10.1287/inte.18.5.55
发表时间: 1988
期刊: Interfaces
影响因子: --
作者:
D. White
通讯作者: D. White
分析和逃避规划中的局​​部最优作为部分可观察域的推理
DOI: 10.1007/978-3-642-23783-6_39
发表时间: 2011
期刊: ArXiv
影响因子: --
作者:
P. Poupart;Tobias Lang;Marc Toussaint
通讯作者: Marc Toussaint
马尔可夫决策过程的几何策略迭代
DOI: 10.1145/3534678.3539478
发表时间: 2022
期刊: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子: --
作者:
Yue Wu;J. D. Loera
通讯作者: J. D. Loera