CMRM: A Cross-Modal Reasoning Model to Enable Zero-Shot Imitation Learning for Robotic RFID Inventory in Unstructured Environments
CMRM: A Cross-Modal Reasoning Model to Enable Zero-Shot Imitation Learning for Robotic RFID Inventory in Unstructured Environments
复制标题
DOI:
10.1109/globecom54140.2023.10437833
复制
发表时间:
2023-12
期刊:
影响因子:
--
通讯作者:
Yongshuai Wu;Jian Zhang;Shaoen Wu;Shiwen Mao;Ying Wang
中科院分区:
文献类型:
--
作者:
Yongshuai Wu;Jian Zhang;Shaoen Wu;Shiwen Mao;Ying Wang
The fast development in Deep Learning (DL) has made it a promising technique for various autonomous robotic systems. Recently, researchers have explored deploying DL models, such as Reinforcement Learning and Imitation Learning, to enable robots for Radio-frequency Identification (RFID) based inventory tasks. However, the existing methods are either focused on a single field or need tremendous data and time to train. To address these problems, this paper presents a Cross-Modal Reasoning Model (CMRM), which is designed to extract high-dimension information from multiple sensors and learn to reason from spatial and historical features for latent cross-modal relations. Furthermore, CMRM aligns the learned tasking policy to high-level features to offer zero-shot generalization to unseen environments. We conduct extensive experiments in several virtual environments as well as in indoor settings with robots for RFID inventory. The experimental results demonstrate that the proposed CMRM can significantly improve learning efficiency by around 20 times. It also demonstrates a robust zero-shot generalization for deploying a learned policy in unseen environments to perform RFID inventory tasks successfully.