Training Agents With Interactive Reinforcement Learning and Contextual Affordances

Training Agents With Interactive Reinforcement Learning and Contextual Affordances
复制标题

DOI:
10.1109/tcds.2016.2543839
复制
发表时间:
2016-12-01
影响因子:
5
通讯作者:
Wermter, Stefan
Wermter, Stefan
中科院分区:
计算机科学3区
文献类型:
--
作者:
Cruz, Francisco;Magg, Sven;Wermter, Stefan

文献摘要

被引文献

相似文献

未来,机器人将被更广泛地用作家庭场景中的助手,并且必须能够通过跨模式互动学习从训练员那里获得专业知识。一种很有前途的方法是交互式强化学习(IRL),即外部培训师就加快学习过程的行动向学徒提供建议。在本文中,我们提出了一种用于家庭清洁桌子任务的IRL方法,并使用模拟机器人比较了三种不同的学习方法:1)强化学习(RL);2)具有上下文负担的RL以避免失败状态;3)先前训练的机器人作为第二个学徒机器人的训练器。然后,我们证明了IRL的使用导致了不同水平的交互和反馈一致性的不同表现。实验结果表明,模拟机器人虽然工作速度慢、成功率低,但仍能在RL的作用下完成任务。有了RL和上下文负担,只需要更少的行动,就可以达到更高的成功率。对于IRL的良好表现,必须考虑反馈的一致性程度,因为不一致性可能会导致学习过程中的相当大的延迟。总体而言,我们证明了在大多数学习情况下,交互反馈为机器人提供了优势。
In the future, robots will be used more extensively as assistants in home scenarios and must be able to acquire expertise from trainers by learning through crossmodal interaction. One promising approach is interactive reinforcement learning (IRL) where an external trainer advises an apprentice on actions to speed up the learning process. In this paper we present an IRL approach for the domestic task of cleaning a table and compare three different learning methods using simulated robots: 1) reinforcement learning (RL); 2) RL with contextual affordances to avoid failed states; and 3) the previously trained robot serving as a trainer to a second apprentice robot. We then demonstrate that the use of IRL leads to different performance with various levels of interaction and consistency of feedback. Our results show that the simulated robot completes the task with RL, although working slowly and with a low rate of success. With RL and contextual affordances fewer actions are needed and can reach higher rates of success. For good performance with IRL it is essential to consider the level of consistency of feedback since inconsistencies can cause considerable delay in the learning process. In general, we demonstrate that interactive feedback provides an advantage for the robot in most of the learning cases.