Rewarding Behaviors

Rewarding Behaviors
复制标题

奖励行为

DOI:
--
复制
发表时间:
1996
期刊:
AAAI/IAAI, Vol. 2
影响因子:
--
通讯作者:
Adam J. Grove
Adam J. Grove
中科院分区:
--
文献类型:
--
作者:
F. Bacchus;Craig Boutilier;Adam J. Grove

文献摘要

被引文献

相似文献

马尔可夫决策过程(MDP)是决策理论计划(DTP)的非常流行的工具,部分是由于包含有效解决方案技术的良好开发,表现力的理论。但是,马尔可夫的假设(动态和奖励仅取决于当前状态,而不是历史)通常是不合适的。奖励尤其如此:我们经常希望将奖励与随着时间的推移延伸的行为联系起来。当然,如果我们拥有足够丰富的状态空间(状态编码足够的历史),则可以在MDP中编码此类奖励过程。但是,通常很难“手工制作”适当的状态空间来编码适当的历史记录。 在非马克维亚奖励是通过将值分配给时间逻辑的公式来编码的情况下,我们认为这个问题。这些公式表征了时间扩展行为的价值。我们认为,这允许自然表示许多常见的非马克维亚奖励。主要结果是一种算法,鉴于以这种方式表达的非马克维亚奖励的决策过程,它会自动构建等效的MDP(具有马尔可夫奖励结构),从而允许使用标准技术进行最佳的政策构建。
Markov decision processes (MDPs) are a very popular tool for decision theoretic planning (DTP), partly because of the welldeveloped, expressive theory that includes effective solution techniques. But the Markov assumption--that dynamics and rewards depend on the current state only, and not on history-- is often inappropriate. This is especially true of rewards: we frequently wish to associate rewards with behaviors that extend over time. Of course, such reward processes can be encoded in an MDP should we have a rich enough state space (where states encode enough history). However it is often difficult to "hand craft" suitable state spaces that encode an appropriate amount of history. We consider this problem in the case where non-Markovian rewards are encoded by assigning values to formulas of a temporal logic. These formulas characterize the value of temporally extended behaviors. We argue that this allows a natural representation of many commonly encountered non-Markovian rewards. The main result is an algorithm which, given a decision process with non-Markovian rewards expressed in this manner, automatically constructs an equivalent MDP (with Markovian reward structure), allowing optimal policy construction using standard techniques.