Learning to Teach Reinforcement Learning Agents

Learning to Teach Reinforcement Learning Agents
复制标题

DOI:
10.3390/make1010002
复制
发表时间:
2019-03-01
影响因子:
3.9
通讯作者:
Vlahavas, Ioannis
Vlahavas, Ioannis
中科院分区:
其他
文献类型:
--
作者:
Fachantidis, Anestis;Taylor, Matthew;Vlahavas, Ioannis

文献摘要

被引文献

相似文献

在这篇文章中,我们研究了预算下行动建议的迁移学习模型。我们专注于强化学习的教师提供行动建议,异构的学生玩吃豆人游戏在有限的建议预算。首先,我们研究了几个关键因素影响咨询质量在这种情况下,如教师的平均表现,其方差和奖励折扣的重要性,在建议。实验表明,最好的表演者并不总是最好的老师,并揭示了非平凡的重要性的变异系数(CV)作为一个统计选择的政策,产生的建议。CV统计量将方差与相应的均值相关联。其次,本文研究了在预算下分配建议的政策学习。鉴于相关文献中的大多数方法都依赖于建议分配的算法,我们将问题制定为学习问题,并提出了一种新的强化学习算法,能够学习何时建议或不建议。所提出的算法是能够建议,即使它不知道学生的预期行动,需要显着减少训练时间相比,以前的学习方法。最后,在本文中,我们认为学习在预算下提供建议是一个更通用的学习问题的例子:约束剥削强化学习。
In this article, we study the transfer learning model of action advice under a budget. We focus on reinforcement learning teachers providing action advice to heterogeneous students playing the game of Pac-Man under a limited advice budget. First, we examine several critical factors affecting advice quality in this setting, such as the average performance of the teacher, its variance and the importance of reward discounting in advising. The experiments show that the best performers are not always the best teachers and reveal the non-trivial importance of the coefficient of variation (CV) as a statistic for choosing policies that generate advice. The CV statistic relates variance to the corresponding mean. Second, the article studies policy learning for distributing advice under a budget. Whereas most methods in the relevant literature rely on heuristics for advice distribution, we formulate the problem as a learning one and propose a novel reinforcement learning algorithm capable of learning when to advise or not. The proposed algorithm is able to advise even when it does not have knowledge of the student's intended action and needs significantly less training time compared to previous learning approaches. Finally, in this article, we argue that learning to advise under a budget is an instance of a more generic learning problem: Constrained Exploitation Reinforcement Learning.