Stochastic optimal control with dynamic, time-consistent risk constraints

Stochastic optimal control with dynamic, time-consistent risk constraints
复制标题

具有动态、时间一致风险约束的随机最优控制

DOI:
--
复制
发表时间:
2013
期刊:
American Control Conference
影响因子:
--
通讯作者:
M. Pavone
M. Pavone
中科院分区:
--
文献类型:
--
作者:
Yinlam Chow;M. Pavone

文献摘要

被引文献

相似文献

在本文中,我们提出了一个动态规划方法的随机最优控制问题的动态,时间一致的风险约束。在过去的20年里,约束随机最优控制问题,当人们必须考虑多个目标时自然出现,已经被广泛研究;然而,在大多数公式中,约束被公式化为风险中性(即,通过考虑预期成本),或者通过应用静态的、单周期风险度量而对“时间一致性”关注有限(即,这些指标是否确保多个时期风险偏好的合理一致性)。最近,重大进展已经取得了严格的理论发展动态,时间一致的风险度量多周期(风险敏感)的决策过程,然而,他们的集成约束随机最优控制问题很少受到关注。本文的目的就是弥合这一差距。首先,我们制定的随机最优控制问题的动态,时间一致的风险约束,我们的特征尾巴子问题(这需要添加一个马尔可夫结构的风险度量)。其次,我们开发了一个动态规划方法的解决方案,它允许计算最优成本的价值迭代。最后,我们提出了一个程序来构建最优策略。
In this paper we present a dynamic programming approach to stochastic optimal control problems with dynamic, time-consistent risk constraints. Constrained stochastic optimal control problems, which naturally arise when one has to consider multiple objectives, have been extensively investigated in the past 20 years; however, in most formulations, the constraints are formulated as either risk-neutral (i.e., by considering an expected cost), or by applying static, single-period risk metrics with limited attention to “time-consistency” (i.e., to whether such metrics ensure rational consistency of risk preferences across multiple periods). Recently, significant strides have been made in the development of a rigorous theory of dynamic, time-consistent risk metrics for multi-period (risk-sensitive) decision processes; however, their integration within constrained stochastic optimal control problems has received little attention. The goal of this paper is to bridge this gap. First, we formulate the stochastic optimal control problem with dynamic, time-consistent risk constraints and we characterize the tail subproblems (which requires the addition of a Markovian structure to the risk metrics). Second, we develop a dynamic programming approach for its solution, which allows to compute the optimal costs by value iteration. Finally, we present a procedure to construct optimal policies.