MUTEX: Learning Unified Policies from Multimodal Task Specifications

MUTEX: Learning Unified Policies from Multimodal Task Specifications
复制标题

DOI:
10.48550/arxiv.2309.14320
复制
发表时间:
2023-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Rutav Shah;Roberto Mart'in-Mart'in-Roberto-Mart'in-Mart'in-2196148773;Yuke Zhu
Rutav Shah;Roberto Mart'in-Mart'in-Roberto-Mart'in-Mart'in-2196148773;Yuke Zhu
中科院分区:
其他
文献类型:
--
作者:
Rutav Shah;Roberto Mart'in-Mart'in-Roberto-Mart'in-Mart'in-2196148773;Yuke Zhu

文献摘要

相似文献

人们使用不同的方式,如语音、文本、图像、视频等,与队友交流他们的意图和目标。为了让机器人成为更好的助手,我们的目标是赋予它们遵循指令和理解人类伙伴指定任务的能力。大多数机器人策略学习方法都只关注任务规范的单一模态,而忽略了丰富的跨模态信息。我们提出了MUTEX,一种从多模态任务规范中学习策略的统一方法。它训练了一个基于变压器的架构来促进跨模态推理,在两个阶段的训练过程中结合了掩模建模和跨模态匹配目标。经过训练,MUTEX可以在六种学习模式(视频演示、目标图像、文本目标描述、文本指令、语音目标描述和语音指令)中的任何一种模式中遵循任务规范,或者它们的组合。我们在一个新设计的数据集中系统地评估了MUTEX的好处,该数据集中有100个模拟任务和50个现实世界中的任务,用不同模式的任务规范的多个实例进行了注释,并观察了针对任何单一模式专门训练的方法的性能改进。更多信息请访问https://ut-austin-rpl.github.io/MUTEX/
Humans use different modalities, such as speech, text, images, videos, etc., to communicate their intent and goals with teammates. For robots to become better assistants, we aim to endow them with the ability to follow instructions and understand tasks specified by their human partners. Most robotic policy learning methods have focused on one single modality of task specification while ignoring the rich cross-modal information. We present MUTEX, a unified approach to policy learning from multimodal task specifications. It trains a transformer-based architecture to facilitate cross-modal reasoning, combining masked modeling and cross-modal matching objectives in a two-stage training procedure. After training, MUTEX can follow a task specification in any of the six learned modalities (video demonstrations, goal images, text goal descriptions, text instructions, speech goal descriptions, and speech instructions) or a combination of them. We systematically evaluate the benefits of MUTEX in a newly designed dataset with 100 tasks in simulation and 50 tasks in the real world, annotated with multiple instances of task specifications in different modalities, and observe improved performance over methods trained specifically for any single modality. More information at https://ut-austin-rpl.github.io/MUTEX/