CORA: Benchmarks, Baselines, and Metrics as a Platform for Continual Reinforcement Learning Agents

CORA: Benchmarks, Baselines, and Metrics as a Platform for Continual Reinforcement Learning Agents
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Sam Powers;Eliot Xing;Eric Kolve;Roozbeh Mottaghi;A. Gupta
Sam Powers;Eliot Xing;Eric Kolve;Roozbeh Mottaghi;A. Gupta
中科院分区:
其他
文献类型:
--
作者:
Sam Powers;Eliot Xing;Eric Kolve;Roozbeh Mottaghi;A. Gupta

文献摘要

相似文献

持续强化学习的进展由于几个进入障碍而受到限制:缺少代码,高计算要求以及缺乏合适的基准。在这项工作中,我们提出了CORA,一个持续强化学习代理的平台,在单个代码包中提供基准,基线和指标。我们提供的基准测试旨在评估持续RL挑战的不同方面,例如灾难性遗忘,可塑性,概括能力和样本有效学习。三个基准测试使用视频游戏环境(Atari,Procgen,NetHack)。第四个基准,CHORES,由四个不同的任务序列在视觉上逼真的家庭模拟器,从一组不同的任务和场景参数。为了在这些基准测试中比较持续强化学习方法,我们在CORA中准备了三个指标:持续评估,孤立遗忘和零触发前向传输。最后,CORA包括一组现有算法的高性能开源基线,供研究人员使用和扩展。我们发布CORA,并希望持续强化学习社区能够从我们的贡献中受益,以加速新的持续强化学习算法的开发。
Progress in continual reinforcement learning has been limited due to several barriers to entry: missing code, high compute requirements, and a lack of suitable benchmarks. In this work, we present CORA, a platform for Continual Reinforcement Learning Agents that provides benchmarks, baselines, and metrics in a single code package. The benchmarks we provide are designed to evaluate different aspects of the continual RL challenge, such as catastrophic forgetting, plasticity, ability to generalize, and sample-efficient learning. Three of the benchmarks utilize video game environments (Atari, Procgen, NetHack). The fourth benchmark, CHORES, consists of four different task sequences in a visually realistic home simulator, drawn from a diverse set of task and scene parameters. To compare continual RL methods on these benchmarks, we prepare three metrics in CORA: Continual Evaluation, Isolated Forgetting, and Zero-Shot Forward Transfer. Finally, CORA includes a set of performant, open-source baselines of existing algorithms for researchers to use and expand on. We release CORA and hope that the continual RL community can benefit from our contributions, to accelerate the development of new continual RL algorithms.