A Multi-Agent Approach for Adaptive Finger Cooperation in Learning-based In-Hand Manipulation

A Multi-Agent Approach for Adaptive Finger Cooperation in Learning-based In-Hand Manipulation
复制标题

DOI:
10.1109/icra48891.2023.10160909
复制
发表时间:
2022-10
期刊:
2023 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Lingfeng Tao;Jiucai Zhang;Michael Bowman;Xiaoli Zhang
Lingfeng Tao;Jiucai Zhang;Michael Bowman;Xiaoli Zhang
中科院分区:
其他
文献类型:
--
作者:
Lingfeng Tao;Jiucai Zhang;Michael Bowman;Xiaoli Zhang

文献摘要

相似文献

由于多指机械手的自由度高,与物体的交互复杂,因此其手内操作具有挑战性。为了实现手内操作,现有的基于深度强化学习的方法主要集中在通过集中式学习机制训练单个机器人结构特定的策略,缺乏对机器人故障等变化的适应性。为了解决这一问题,本文将每个手指视为一个独立的智能体,训练多个智能体控制各自指定的手指协同完成手内操作任务。我们提出了多智能体全局观察批评和局部观察行动者(MAGCLA)的方法,其中的评论家可以观察所有的代理人的行动全球,和行动者只在本地观察其邻居的行动。此外,传统的个人经验重放可能会导致不稳定的合作,由于每个代理的异步性能增量,这是关键的在手操作任务。为了解决这个问题,我们提出了同步后见之明的经验重放(SHER)方法,同步和有效地重用所有代理的重放经验。在两个在手操作任务的影子灵巧手的方法进行评估。结果表明,SHER有助于MAGCLA实现可比的学习效率,一个单一的政策,和MAGCLA的方法是在不同的任务更具有推广性。训练后的策略具有更高的适应性,在机器人故障测试相比,基线多智能体和单智能体的方法。
In-hand manipulation is challenging for a multi-finger robotic hand due to its high degrees of freedom and complex interaction with the object. To enable in-hand manipulation, existing deep reinforcement learning-based approaches mainly focus on training a single robot-structure-specific policy through the centralized learning mechanism, lacking adaptability to changes like robot malfunction. To solve this limitation, this work treats each finger as an individual agent and trains multiple agents to control their assigned fingers to complete the in-hand manipulation task cooperatively. We propose the Multi-Agent Global-Observation Critic and Local-Observation Actor (MAGCLA) method, where the critic can observe all agents' actions globally, and the actor only locally observes its neighbors' actions. Besides, conventional individual experience replay may cause unstable cooperation due to the asynchronous performance increment of each agent, which is critical for in-hand manipulation tasks. To solve this issue, we propose the Synchronized Hindsight Experience Replay (SHER) method to synchronize and efficiently reuse the replayed experience across all agents. The methods are evaluated in two in-hand manipulation tasks on the Shadow dexterous hand. The results show that SHER helps MAGCLA achieve comparable learning efficiency to a single policy, and the MAGCLA approach is more generalizable in different tasks. The trained policies have higher adaptability in the robot malfunction test compared to the baseline multi-agent and single-agent approaches.