Steady-State Planning in Expected Reward Multichain MDPs

Steady-State Planning in Expected Reward Multichain MDPs
复制标题

DOI:
10.1613/jair.1.12611
复制
发表时间:
2020-12
期刊:
J. Artif. Intell. Res.
影响因子:
--
通讯作者:
George K. Atia;Andre Beckus;Ismail R. Alkhouri;Alvaro Velasquez
George K. Atia;Andre Beckus;Ismail R. Alkhouri;Alvaro Velasquez
中科院分区:
其他
文献类型:
--
作者:
George K. Atia;Andre Beckus;Ismail R. Alkhouri;Alvaro Velasquez

文献摘要

被引文献

相似文献

规划领域对决策政策的正式综合越来越感兴趣。这种形式综合通常需要找到一种满足某种明确定义逻辑形式的形式规范的策略。尽管许多此类逻辑在捕获期望的代理行为的能力方面具有不同程度的表达性和复杂性,但在推导满足一般系统模型中某些类型的渐近行为的决策策略时,它们的价值是有限的。特别是,我们对指定代理的稳态行为的约束感兴趣,它捕获代理在与环境无限期交互时在每个状态中花费的时间比例。这有时被称为代理的平均或预期行为,并且相关的规划问题面临着重大挑战,除非在其图结构的连接性方面对底层模型施加严格的限制。在本文中,我们探讨了这个稳态规划问题,其中包括为代理导出决策策略,以满足其稳态行为的约束。提出了针对多链马尔可夫决策过程(MDP)一般情况的线性规划解决方案,并且我们证明了所提出的程序的最优解决方案产生具有严格行为保证的固定策略。
The planning domain has experienced increased interest in the formal synthesis of decision-making policies. This formal synthesis typically entails finding a policy which satisfies formal specifications in the form of some well-defined logic. While many such logics have been proposed with varying degrees of expressiveness and complexity in their capacity to capture desirable agent behavior, their value is limited when deriving decision-making policies which satisfy certain types of asymptotic behavior in general system models. In particular, we are interested in specifying constraints on the steady-state behavior of an agent, which captures the proportion of time an agent spends in each state as it interacts for an indefinite period of time with its environment. This is sometimes called the average or expected behavior of the agent and the associated planning problem is faced with significant challenges unless strong restrictions are imposed on the underlying model in terms of the connectivity of its graph structure. In this paper, we explore this steady-state planning problem that consists of deriving a decision-making policy for an agent such that constraints on its steady-state behavior are satisfied. A linear programming solution for the general case of multichain Markov Decision Processes (MDPs) is proposed and we prove that optimal solutions to the proposed programs yield stationary policies with rigorous guarantees of behavior.