Planning for human-robot interaction: representing time and human intention

Planning for human-robot interaction: representing time and human intention
复制标题

人机交互规划:代表时间和人类意图

DOI:
--
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
F. Broz
F. Broz
中科院分区:
--
文献类型:
--
作者:
I. Nourbakhsh;R. Simmons;F. Broz

文献摘要

参考文献

被引文献

相似文献

本文提出了一种新的方法来规划一类特定的人-机器人交互领域:在这些领域中,机器人与人类一起从事受社会习俗支配的任务。当人类执行这些任务时,他们试图在与他人共享的环境中实现个人目标。社会习俗的存在是作为如何与他人互动的指导方针,以便有关各方能够有效地实现他们的目标,而不相互干扰。认识到他人正在努力实现的目标,并在互动中的适当时间采取行动,这是社交能力的关键能力。 本文所采用的人-机器人社会交互方法致力于为规划建立更准确的社会任务模型。由于人类参与者被建模为环境的一部分,因此这些问题中的世界状态是动态的和部分可观测的。在部分可观测马尔可夫决策过程(POMDP)中,将人的意图表示为隐藏状态,并对行为结果的时间相关性进行显式建模。由人类专家设计的模型结构与人类任务性能数据相结合。由此产生的模型既大又复杂。在状态空间的时间维度上的状态聚集被用来在表示的准确性和其大小之间进行权衡,以便找到也可以容易地求解的具有足够表现力的模型。这种方法的实用性通过实现移动机器人的控制器和驾驶模拟器中的代理来演示,该控制器可以与人一起乘坐电梯,驾驶模拟器可以执行匹兹堡留给人类司机的任务。通过将使用建议的建模技术获得的策略与使用表达能力较低的表示法开发的策略进行比较来评估性能。在与人类参与者的交互中,以人类意图为隐含状态的时变POMDP模型的策略优于其他策略,在自然性和社会合适性方面获得了更高的回报和更积极的评价。
This thesis proposes a novel approach to planning for a specific class of human-robot interaction domains: those in which robots engage in tasks with humans that are governed by social conventions. When humans perform these tasks, they try to achieve individual goals in an environment that they share with other people. Social conventions exist as a guideline for how to interact with others so that all parties involved can achieve their goals efficiently without interfering with one another. Recognizing what goals others are trying to achieve and performing actions at the appropriate time in the interaction are critical abilities for social competence. The approach to human-robot social interaction taken in this thesis focuses on creating more accurate models of social tasks for planning. Because the human participants are modeled as a part of the environment, the world state in these problems is dynamic and partially observable. Human intention is represented as hidden state in a partially observable Markov decision process (POMDP), and the time-dependence of action outcomes are explicitly modeled. A model structure designed by a human expert is combined with human task performance data. The resulting models are large and complex. State aggregation over the time dimension of the state space is used to trade off between the accuracy of the representation and its size in order to find sufficiently expressive models that can also be solved tractably. The utility of this approach is demonstrated by implementing a controller for a mobile robot that rides elevators with people and an agent in a driving simulator that performs the Pittsburgh left with human drivers. Performance is evaluated by comparing the policies obtained using the proposed modeling technique to policies developed using less expressive representations. In an interactions with human participants, the policies for time-dependent POMDP models with human intention as hidden state outperform the other policies, achieving both higher rewards and more positive evaluations for naturalness and social propriety of behavior.
DOI: 10.1037/0012-1649.31.5.838
发表时间: 1995-09-01
影响因子: 4
作者:
MELTZOFF, AN
通讯作者: MELTZOFF, AN