Interactive task learning via embodied corrective feedback

Interactive task learning via embodied corrective feedback
复制标题

通过具体的纠正反馈进行交互式任务学习

DOI:
--
复制
发表时间:
2020
影响因子:
1.9
通讯作者:
A. Lascarides
A. Lascarides
中科院分区:
计算机科学4区
文献类型:
--
作者:
Mattias Appelgren;A. Lascarides

文献摘要

参考文献

被引文献

相似文献

本文讨论了交互式任务学习中的一项任务(Laird 等人。IEEE Intelll Syst 32:6–21, 2017)。智能体必须学会建造受规则约束的塔,每当智能体执行违反规则的动作时,老师都会提供口头纠正反馈:例如“不,红色方块应该在蓝色方块上”。智能体必须学会根据这些修正和给出的上下文来构建符合规则的塔。代理不仅在学习过程开始时不了解规则,而且它的领域模型也有缺陷,缺乏表达规则的概念。因此,利用语言证据的智能体必须学习新词的外延,并调整其规划领域的概念化以纳入这些外延。我们表明,通过将话语连贯性对解释的约束纳入学习模型中(Hobbs 着《论话语的连贯性和结构》,斯坦福大学,斯坦福,1985 年;Asher 等人,《对话逻辑》,剑桥大学出版社,剑桥,2003 年),利用语言证据的智能体的表现优于不使用语言证据的强基线。
This paper addresses a task in Interactive Task Learning (Laird et al. IEEE Intell Syst 32:6–21, 2017). The agent must learn to build towers which are constrained by rules, and whenever the agent performs an action which violates a rule the teacher provides verbal corrective feedback: e.g. “No, red blocks should be on blue blocks”. The agent must learn to build rule compliant towers from these corrections and the context in which they were given. The agent is not only ignorant of the rules at the start of the learning process, but it also has a deficient domain model, which lacks the concepts in which the rules are expressed. Therefore an agent that takes advantage of the linguistic evidence must learn the denotations of neologisms and adapt its conceptualisation of the planning domain to incorporate those denotations. We show that by incorporating constraints on interpretation that are imposed by discourse coherence into the models for learning (Hobbs in On the coherence and structure of discourse, Stanford University, Stanford, 1985; Asher et al. in Logics of conversation, Cambridge University Press, Cambridge, 2003), an agent which utilizes linguistic evidence outperforms a strong baseline which does not.
通过从纠正中获取扎根的语言意义来制定学习计划
DOI: --
发表时间: 2019
期刊: --
影响因子: --
作者:
Appelgren M
通讯作者: Appelgren M
从示范中学习的可解释的潜在空间
DOI: --
发表时间: 2018
期刊: --
影响因子: --
作者:
Hristov YS
通讯作者: Hristov YS
连贯性、符号基础和交互式任务学习
DOI: --
发表时间: 2019
期刊: --
影响因子: --
作者:
Appelgren M
通讯作者: Appelgren M