Design for an Optimal Probe

Design for an Optimal Probe
复制标题

最佳探针设计

DOI:
--
复制
发表时间:
2003
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
M. Duff
M. Duff
中科院分区:
--
文献类型:
--
作者:
M. Duff

文献摘要

被引文献

相似文献

给定一个马尔可夫决策过程(MDP),在过程转移概率中表示先验不确定性,我们考虑计算一个优化期望总(有限视界)奖励的策略问题。含蓄地说,这样的策略将有效地解决“探索与开发的权衡”所面临的问题,例如,当一个智能体寻求在与不确定世界的整个互动过程中获得最佳的总强化时。贝叶斯公式导致相关的MDP定义在一组广义过程“超状态”上,其基数随着规划范围呈指数级增长。在这里,我们保留了完整的贝叶斯框架,但通过应用强化学习理论的技术来回避难处。我们将生成的参与者-评论家算法应用于“最优探测”问题,其中的任务是使用在线体验识别MDP的未知转移概率。
Given a Markov decision process (MDP) with expressed prior uncertainties in the process transition probabilities, we consider the problem of computing a policy that optimizes expected total (finite-horizon) reward. Implicitly, such a policy would effectively resolve the "exploration-versus-exploitation tradeoff" faced, for example, by an agent that seeks to optimize total reinforcement obtained over the entire duration of its interaction with an uncertain world. A Bayesian formulation leads to an associated MDP defined over a set of generalized process "hyperstates" whose cardinality grows exponentiaily with the planning horizon. Here we retain the full Bayesian framework, but sidestep intractability by applying techniques from reinforcement learning theory. We apply our resulting actor-critic algorithm to a problem of "optimal probing," in which the task is to identify unknown transition probabilities of an MDP using online experience.