From Play to Policy: Conditional Behavior Generation from Uncurated Robot Data

From Play to Policy: Conditional Behavior Generation from Uncurated Robot Data
复制标题

DOI:
10.48550/arxiv.2210.10047
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Zichen Jeff Cui;Yibin Wang;Nur Muhammad (Mahi) Shafiullah;Lerrel Pinto
Zichen Jeff Cui;Yibin Wang;Nur Muhammad (Mahi) Shafiullah;Lerrel Pinto
中科院分区:
其他
文献类型:
--
作者:
Zichen Jeff Cui;Yibin Wang;Nur Muhammad (Mahi) Shafiullah;Lerrel Pinto

文献摘要

被引文献

相似文献

尽管基于离线数据的大规模序列建模在自然语言和图像生成方面带来了令人印象深刻的性能提升,但将这些想法直接转化为机器人技术一直具有挑战性。造成这种情况的一个关键原因是,从非专业的人类演示者那里收集的未经整理的机器人演示数据(即游戏数据)通常是嘈杂的、多样化的、分布多模态的。这使得从这些数据中提取有用的、以任务为中心的行为成为一个困难的生成建模问题。在这项工作中,我们提出了条件行为转换器(C-BeT),一种将行为转换器的多模态生成能力与未来条件目标规范相结合的方法。在一组模拟基准任务中,我们发现C-BeT在从游戏数据中学习方面平均提高了45.7%。此外,我们首次证明,在没有任何任务标签或奖励信息的情况下,可以在真实世界的机器人上纯粹从游戏数据中学习有用的以任务为中心的行为。机器人视频最好在我们的项目网站上观看:https://play-to-policy.github.io
While large-scale sequence modeling from offline data has led to impressive performance gains in natural language and image generation, directly translating such ideas to robotics has been challenging. One critical reason for this is that uncurated robot demonstration data, i.e. play data, collected from non-expert human demonstrators are often noisy, diverse, and distributionally multi-modal. This makes extracting useful, task-centric behaviors from such data a difficult generative modeling problem. In this work, we present Conditional Behavior Transformers (C-BeT), a method that combines the multi-modal generation ability of Behavior Transformer with future-conditioned goal specification. On a suite of simulated benchmark tasks, we find that C-BeT improves upon prior state-of-the-art work in learning from play data by an average of 45.7%. Further, we demonstrate for the first time that useful task-centric behaviors can be learned on a real-world robot purely from play data without any task labels or reward information. Robot videos are best viewed on our project website: https://play-to-policy.github.io