Act as You Learn: Adaptive Decision-Making in Non-Stationary Markov Decision Processes

Act as You Learn: Adaptive Decision-Making in Non-Stationary Markov Decision Processes
复制标题

DOI:
10.48550/arxiv.2401.01841
复制
发表时间:
2024-01
期刊:
--
影响因子:
--
通讯作者:
Baiting Luo;Yunuo Zhang;Abhishek Dubey;Ayan Mukhopadhyay
Baiting Luo;Yunuo Zhang;Abhishek Dubey;Ayan Mukhopadhyay
中科院分区:
其他
文献类型:
--
作者:
Baiting Luo;Yunuo Zhang;Abhishek Dubey;Ayan Mukhopadhyay

文献摘要

相似文献

顺序决策中的一个基本(而且很大程度上是公开的)挑战是处理非平稳环境,即外部环境条件随时间而变化。这些问题传统上被建模为非平稳马尔可夫决策过程(NSMDP)。然而,现有的NSMDP决策方法有两个主要缺点:首先,它们假设当前更新的环境动态是已知的(尽管未来动态可能会改变);其次,规划在很大程度上是悲观的,即,代理人“安全地”采取行动,以说明环境的非平稳演变。我们认为,这两个假设是无效的,在实践中-更新的环境条件是很少知道的,作为代理与环境的互动,它可以了解更新的动态和避免悲观,至少在国家的动态是有信心的。我们提出了一种启发式搜索算法,称为\textit{自适应蒙特卡罗树搜索(ADA-MCTS)},解决这些挑战。我们表明,智能体可以随着时间的推移学习环境的更新动态,然后按照它的学习行为,即,如果代理处于状态空间的一个区域中,它已经更新了关于该区域的知识,则它可以避免悲观。为了量化“更新的知识”,我们分解的任意性和认识的不确定性在代理的更新的信念,并显示代理如何可以使用这些估计决策。我们将所提出的方法与多个成熟的开源问题的决策中的多个最先进的方法进行比较,并根据经验表明,我们的方法更快,适应性更强,而不会牺牲安全性。
A fundamental (and largely open) challenge in sequential decision-making is dealing with non-stationary environments, where exogenous environmental conditions change over time. Such problems are traditionally modeled as non-stationary Markov decision processes (NSMDP). However, existing approaches for decision-making in NSMDPs have two major shortcomings: first, they assume that the updated environmental dynamics at the current time are known (although future dynamics can change); and second, planning is largely pessimistic, i.e., the agent acts ``safely'' to account for the non-stationary evolution of the environment. We argue that both these assumptions are invalid in practice -- updated environmental conditions are rarely known, and as the agent interacts with the environment, it can learn about the updated dynamics and avoid being pessimistic, at least in states whose dynamics it is confident about. We present a heuristic search algorithm called \textit{Adaptive Monte Carlo Tree Search (ADA-MCTS)} that addresses these challenges. We show that the agent can learn the updated dynamics of the environment over time and then act as it learns, i.e., if the agent is in a region of the state space about which it has updated knowledge, it can avoid being pessimistic. To quantify ``updated knowledge,'' we disintegrate the aleatoric and epistemic uncertainty in the agent's updated belief and show how the agent can use these estimates for decision-making. We compare the proposed approach with the multiple state-of-the-art approaches in decision-making across multiple well-established open-source problems and empirically show that our approach is faster and highly adaptive without sacrificing safety.