Multimodal reinforcement learning for partner specific adaptation in robot-multi-robot interaction

Multimodal reinforcement learning for partner specific adaptation in robot-multi-robot interaction
复制标题

机器人与多机器人交互中伙伴特定适应的多模态强化学习

DOI:
10.1109/humanoids53995.2022.10000205
复制
发表时间:
2022
期刊:
IEEE Proceedings on Humanoids 2022, Ginowan, Japan
影响因子:
--
通讯作者:
Oztop Erhan
Oztop Erhan
中科院分区:
--
文献类型:
--
作者:
Kirtay Murat;Hafner Verena V.;Asada Minoru;Kuhlen Anna K.;Oztop Erhan

文献摘要

参考文献

相似文献

成功和高效的团队合作需要了解每个团队成员的专业知识。这些知识通常是在社会交往中获得的,并形成了社会智能,合作伙伴适应行为的基础。本研究的目的是实现这种能力在多个人形机器人的团队。为此,一个人形机器人Nao与三个Pepper机器人进行交互,以执行一个需要整合多模态信息的顺序视听模式回忆任务。Nao将其决策外包(即,动作选择),以通过应用强化学习在神经计算成本方面有效地执行任务。在互动过程中,Nao了解了其合作伙伴的特定专业知识,这使得Nao能够向具有与当前任务状态相对应的专业知识的合作伙伴寻求指导。Nao的认知处理包括多模态自联想记忆,其允许确定感知处理的成本(即,认知负荷)处理视听刺激时。反过来,处理成本由内部奖励生成模块转换为奖励信号。在这种设置中,学习机器人Nao的目标是通过转向其专业知识对应于给定任务状态的伙伴来最小化认知负荷。总的来说,结果表明,学习机器人发现的专业知识的合作伙伴,并利用这些信息来执行其任务,低神经计算成本或认知负荷。
Successful and efficient teamwork requires knowledge of the individual team members' expertise. Such knowledge is typically acquired in social interaction and forms the basis for socially intelligent, partner-adapted behavior. This study aims to implement this ability in teams of multiple humanoid robots. To this end, a humanoid robot, Nao, interacted with three Pepper robots to perform a sequential audio-visual pattern recall task that required integrating multimodal information. Nao outsourced its decisions (i.e., action selections) to its robot partners to perform the task efficiently in terms of neural computational cost by applying reinforcement learning. During the interaction, Nao learned its partners' specific expertise, which allowed Nao to turn for guidance to the partner who has the expertise corresponding to the current task state. The cognitive processing of Nao included a multimodal auto-associative memory that allowed the determination of the cost of perceptual processing (i.e., cognitive load) when processing audio-visual stimuli. In turn, the processing cost is converted into a reward signal by an internal reward generation module. In this setting, the learner robot Nao aims to minimize cognitive load by turning to the partner whose expertise corresponds to a given task state. Overall, the results indicate that the learner robot discovers the expertise of partners and exploits this information to execute its task with low neural computational cost or cognitive load.
qiBullet,用于 Pepper 和 NAO 机器人的基于 Bullet 的模拟器
DOI: --
发表时间: 2019
期刊: arXiv.org
影响因子: --
作者:
Maxime Busy;Maxime Caniot
通讯作者: Maxime Caniot
DOI: --
发表时间: 2019
期刊: Philosophical Transactions of the Royal Society of London. Biological Sciences
影响因子: --
作者:
Samuele Vinanzi;Massimiliano Patacchiola;A. Chella;A. Cangelosi
通讯作者: A. Cangelosi
你会帮助一个悲伤的机器人吗?机器人的情绪表达对人机协作的影响
DOI: --
发表时间: 2020
期刊: IEEE International Symposium on Robot and Human Interactive Communication
影响因子: --
作者:
Shujie Zhou;Leimin Tian
通讯作者: Leimin Tian
基于交互任务中突发情绪的机器人信任建模
DOI: 10.1109/icdl49984.2021.9515645
发表时间: 2021
期刊: 2021 IEEE International Conference on Development and Learning (ICDL)
影响因子: --
作者:
M. Kirtay;Erhan Öztop;M. Asada;V. Hafner
通讯作者: V. Hafner
DOI: 10.3389/frobt.2018.00075
发表时间: 2018
影响因子: 3.4
作者:
Winfield AFT
通讯作者: Winfield AFT