Expedited Learning in MDPs with Side Information

Expedited Learning in MDPs with Side Information
复制标题

DOI:
10.1109/cdc.2018.8619134
复制
发表时间:
2018-12
期刊:
2018 IEEE Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Melkior Ornik;Jie Fu;Niklas T. Lauffer;W. K. Perera;Mohammed Alshiekh;M. Ono;U. Topcu
Melkior Ornik;Jie Fu;Niklas T. Lauffer;W. K. Perera;Mohammed Alshiekh;M. Ono;U. Topcu
中科院分区:
其他
文献类型:
--
作者:
Melkior Ornik;Jie Fu;Niklas T. Lauffer;W. K. Perera;Mohammed Alshiekh;M. Ono;U. Topcu

文献摘要

相似文献

在转移概率未知的马尔可夫决策过程中,控制策略的标准合成方法主要依赖于探索和开发的结合。虽然这些方法通常提供理论上的保证系统的性能,所需的时间步长和样本的数量,以初步探索的环境之前,合成一个性能良好的控制策略是不切实际的大。本文通过将先验的现有知识纳入学习中,当这些知识可用时,部分地消除了这样的负担。基于关于不同状态下的转移概率之间的差异的界限的先验信息,我们提出了一种学习方法,其中在给定状态下的转移概率不仅从在该状态下重复执行某个动作的结果中学习,而且还从在已知具有相似转移概率的状态下执行动作的结果中学习。由于直接获得的信息在确定转移概率时比二手信息更可靠,即,从类似但可能略有不同的状态获得的信息,间接获得的样本相对于转移概率差的已知界限进行加权。虽然所提出的策略可以自然地导致学习的转移概率的错误,我们表明,通过适当的选择的权重,这样的错误可以减少,形成一个接近最优的控制策略在贝叶斯意义上所需的步骤的数量可以显着减少。
Standard methods for synthesis of control policies in Markov decision processes with unknown transition probabilities largely rely on a combination of exploration and exploitation. While these methods often offer theoretical guarantees on system performance, the number of time steps and samples needed to initially explore the environment before synthesizing a well-performing control policy is impractically large. This paper partially alleviates such a burden by incorporating a priori existing knowledge into learning, when such knowledge is available. Based on prior information about bounds on the differences between the transition probabilities at different states, we propose a learning approach where the transition probabilities at a given state are not only learned from outcomes of repeatedly performing a certain action at that state, but also from outcomes of performing actions at states that are known to have similar transition probabilities. Since the directly obtained information is more reliable at determining transition probabilities than second-hand information, i.e., information obtained from similar but potentially slightly different states, samples obtained indirectly are weighted with respect to the known bounds on the differences of transition probabilities. While the proposed strategy can naturally lead to errors in learned transition probabilities, we show that, by proper choice of the weights, such errors can be reduced, and the number of steps needed to form a near-optimal control policy in the Bayesian sense can be significantly decreased.