Modeling robot trust based on emergent emotion in an interactive task

Modeling robot trust based on emergent emotion in an interactive task
复制标题

基于交互任务中突发情绪的机器人信任建模

DOI:
10.1109/icdl49984.2021.9515645
复制
发表时间:
2021
期刊:
2021 IEEE International Conference on Development and Learning (ICDL)
影响因子:
--
通讯作者:
V. Hafner
V. Hafner
中科院分区:
--
文献类型:
--
作者:
M. Kirtay;Erhan Öztop;M. Asada;V. Hafner

文献摘要

被引文献

相似文献

信任是人与人、人与机器人互动的重要组成部分。在这些相互作用中发挥重要作用的因素一直是机器人技术中一个有吸引力的问题。然而,旨在开发机器人对交互伙伴信任的计算模型的研究仍然相对有限。在这项研究中,我们扩展了我们的涌现情感模型,提出机器人对交互伙伴(即受托人)的信任可以通过交互对机器人(即信任者)计算能量预算的影响来建立。具体来说,我们展示了如何通过决策框架中感知处理(例如,用于视觉回忆的视觉刺激处理)的计算成本来建模代理的高级情绪(例如,幸福感)。为了实现这种方法,我们赋予 Pepper 人形机器人两个模块:一个自动联想存储器,用于提取执行视觉回忆所需的计算能量;以及一个内部奖励机制,用于指导无模型强化学习以产生计算能量成本感知行为。通过这种设置,机器人可以使用不同的指导策略(即可靠、不太可靠和随机)与在线教练进行交互。通过与导师的交互,机器人根据感知处理成本关联累积奖励值,以评估导师并确定应该信任哪位导师。总的来说,结果表明机器人可以区分教练的指导策略。此外,在自由选择的情况下,机器人会信任能够增加总奖励的可靠选择,从而减少执行下一项任务所需的计算能量(认知负荷)。
Trust is an essential component in human-human and human-robot interactions. The factors that play potent roles in these interactions have been an attractive issue in robotics. However, the studies that aim at developing a computational model of robot trust in interaction partners remain relatively limited. In this study, we extend our emergent emotion model to propose that the robot’s trust in the interaction partner (i.e., trustee) can be established by the effect of the interactions on the computational energy budget of the robot (i.e., trustor). To be concrete, we show how high-level emotions (e.g., wellbeing) of an agent can be modeled by the computational cost of perceptual processing (e.g., visual stimulus processing for visual recalling) in a decision-making framework. To realize this approach, we endow the Pepper humanoid robot with two modules: an auto-associative memory that extracts the required computational energy to perform a visual recalling, and an internal reward mechanism guiding model-free reinforcement learning to yield computational energy cost-aware behaviors. With this setup, the robot interacts with online instructors with different guiding strategies, namely reliable, less reliable, and random. Through interaction with the instructors, the robot associates the cumulative reward values based on the cost of perceptual processing to evaluate the instructors and determine which one should be trusted. Overall the results indicate that the robot can differentiate the guiding strategies of the instructors. Additionally, in the case of free choice, the robot trusts the reliable one that increases the total reward – and therefore reduces the required computational energy (cognitive load)– to perform the next task.