Leveraging Observations in Bandits: Between Risks and Benefits

Leveraging Observations in Bandits: Between Risks and Benefits
复制标题

利用对强盗的观察:风险与收益之间

DOI:
--
复制
发表时间:
2019
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Doina Precup
Doina Precup
中科院分区:
--
文献类型:
--
作者:
A. Lupu;A. Durand;Doina Precup

文献摘要

被引文献

相似文献

模仿学习已被广泛用于加速新手代理的学习,允许他们利用专家的现有数据。允许代理受到外部观察的影响可以有利于学习过程,但它也会使代理面临以下次优行为的风险。在本文中,我们研究这个问题的背景下,土匪。更具体地说,我们认为一个代理(学习者)是一个土匪式的决策任务进行交互,但也可以观察到与同一环境的目标政策。学习者只观察目标的行为,而不是获得的奖励。我们引入了一个新的强盗乐观修改器,使用条件乐观视乎行动的目标,以指导代理的探索。我们分析了这种修改的效果上著名的置信上限算法证明,它保留了遗憾的上限为O(lnT),即使在存在一个非常差的目标,我们得出的依赖性的预期遗憾的一般目标政策。我们提供的实证结果显示了巨大的好处,以及某些固有的局限性,观察学习在多臂土匪设置。采用大概率满足理论假设的目标进行实验,缩小了理论与应用之间的差距。
Imitation learning has been widely used to speed up learning in novice agents, by allowing them to leverage existing data from experts. Allowing an agent to be influenced by external observations can benefit to the learning process, but it also puts the agent at risk of following sub-optimal behaviours. In this paper, we study this problem in the context of bandits. More specifically, we consider that an agent (learner) is interacting with a bandit-style decision task, but can also observe a target policy interacting with the same environment. The learner observes only the target’s actions, not the rewards obtained. We introduce a new bandit optimism modifier that uses conditional optimism contingent on the actions of the target in order to guide the agent’s exploration. We analyze the effect of this modification on the well-known Upper Confidence Bound algorithm by proving that it preserves a regret upper-bound of order O(lnT), even in the presence of a very poor target, and we derive the dependency of the expected regret on the general target policy. We provide empirical results showing both great benefits as well as certain limitations inherent to observational learning in the multi-armed bandit setting. Experiments are conducted using targets satisfying theoretical assumptions with high probability, thus narrowing the gap between theory and application.