Learning From Explanations Using Sentiment and Advice in RL

Learning From Explanations Using Sentiment and Advice in RL
复制标题

在强化学习中使用情感和建议从解释中学习

DOI:
10.1109/tcds.2016.2628365
复制
发表时间:
2017
影响因子:
5
通讯作者:
A. Thomaz
A. Thomaz
中科院分区:
计算机科学3区
文献类型:
--
作者:
Samantha Krening;Brent Harrison;K. Feigh;C. Isbell;Mark O. Riedl;A. Thomaz

文献摘要

被引文献

相似文献

为了让机器人向没有机器学习专业知识的人学习,机器人应该从自然的人类指令中学习。大多数包含解释的机器学习技术要求人们使用有限的词汇表并提供状态信息,即使它并不直观。本文讨论了一个软件代理,学会了玩马里奥兄弟游戏使用的解释。我们改善从解释中学习的目标有两个:1)将解释过滤为建议和警告,2)从没有状态信息的句子中学习政策。我们使用情感分析来过滤解释,以提供做什么的建议和避免什么的警告。我们开发了以对象为中心的建议,以表示代理在处理对象时应该采取什么行动。强化学习代理使用以对象为中心的建议来学习最大化其奖励的策略。在减少假阴性之后,使用情感作为过滤器的准确率约为85%。以对象为中心的建议比没有建议时表现得更好,代理人知道在哪里应用建议,代理人可以从对抗性建议中恢复过来。我们还发现,互动的方法应该被设计为减轻人类教师的认知负荷,否则建议的质量可能会很差。
In order for robots to learn from people with no machine learning expertise, robots should learn from natural human instruction. Most machine learning techniques that incorporate explanations require people to use a limited vocabulary and provide state information, even if it is not intuitive. This paper discusses a software agent that learned to play the Mario Bros. game using explanations. Our goals to improve learning from explanations were twofold: 1) to filter explanations into advice and warnings and 2) to learn policies from sentences without state information. We used sentiment analysis to filter explanations into advice of what to do and warnings of what to avoid. We developed object-focused advice to represent what actions the agent should take when dealing with objects. A reinforcement learning agent used object-focused advice to learn policies that maximized its reward. After mitigating false negatives, using sentiment as a filter was approximately 85% accurate. object-focused advice performed better than when no advice was given, the agent learned where to apply the advice, and the agent could recover from adversarial advice. We also found the method of interaction should be designed to ease the cognitive load of the human teacher or the advice may be of poor quality.