Active World Model Learning in Agent-rich Environments with Progress Curiosity

Active World Model Learning in Agent-rich Environments with Progress Curiosity
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
--
影响因子:
--
通讯作者:
Kuno Kim;Megumi Sano;Julian De Freitas;Nick Haber;Daniel L. K. Yamins
Kuno Kim;Megumi Sano;Julian De Freitas;Nick Haber;Daniel L. K. Yamins
中科院分区:
其他
文献类型:
--
作者:
Kuno Kim;Megumi Sano;Julian De Freitas;Nick Haber;Daniel L. K. Yamins

文献摘要

相似文献

世界模型是关于世界如何演变的自我监督预测模型。人类通过好奇地探索他们的环境来学习世界模型,在这个过程中获得高带宽感官输入的紧凑抽象,跨越长时间视野的计划能力,以及对其他代理行为模式的理解。在这项工作中,我们研究了如何设计这样一个好奇心驱动的主动世界模型学习(AWML)系统。为此,我们构建了一个好奇的智能体构建世界模型,同时在视觉上探索一个富含代表性现实世界智能体精华的3D物理环境。我们提出了一个由“-Progress”驱动的AWML系统:一个可扩展且有效的基于学习进度的好奇心信号,并表明“-Progress”自然会产生一个探索策略,该策略以平衡的方式将注意力引向复杂但可学习的动态,从而克服了“白噪声问题”。因此,我们的进度驱动控制器比配备了最先进的探索策略(如随机网络蒸馏和模型不一致)的基线控制器实现了显著更高的AWML性能。
World models are self-supervised predictive models of how the world evolves. Humans learn world models by curiously exploring their environment, in the process acquiring compact abstractions of high bandwidth sensory inputs, the ability to plan across long temporal horizons, and an understanding of the behavioral patterns of other agents. In this work, we study how to design such a curiosity-driven Active World Model Learning (AWML) system. To do so, we construct a curious agent building world models while visually exploring a 3D physical environment rich with distillations of representative real-world agents. We propose an AWML system driven by � -Progress: a scalable and effective learning progress-based curiosity signal and show that � -Progress naturally gives rise to an exploration policy that directs attention to complex but learnable dynamics in a balanced manner, as a result overcoming the “white noise problem”. As a result, our � -Progress-driven controller achieves significantly higher AWML performance than baseline controllers equipped with state-of-the-art exploration strategies such as Random Network Distillation and Model Disagree-ment.