Bounding the Optimal Value Function in Compositional Reinforcement Learning

Bounding the Optimal Value Function in Compositional Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2303.02557
复制
发表时间:
2023-03
期刊:
--
影响因子:
--
通讯作者:
Jacob Adamczyk;Volodymyr Makarenko;A. Arriojas;Stas Tiomkin;R. Kulkarni
Jacob Adamczyk;Volodymyr Makarenko;A. Arriojas;Stas Tiomkin;R. Kulkarni
中科院分区:
其他
文献类型:
--
作者:
Jacob Adamczyk;Volodymyr Makarenko;A. Arriojas;Stas Tiomkin;R. Kulkarni

文献摘要

相似文献

在强化学习(RL)领域,智能体通常负责解决各种问题,只是奖励函数不同。为了用新的奖励函数快速获得未知问题的解决方案,一种流行的方法涉及先前解决的任务的功能组合。然而,之前使用此类函数合成的工作主要集中在合成函数的特定实例上,其限制性假设允许精确的零次合成。我们的工作统一了这些示例,并为标准和熵正则化RL中的组合性提供了一个更一般的框架。我们发现,对于广泛的一类功能,复合任务的最佳解决方案可以与已知的原始任务的解决方案。具体来说,我们提出了双边不等式的最佳复合价值函数的原始任务的价值函数。我们还表明,使用零杆政策的遗憾,可以为这类功能的界。导出的边界可用于开发裁剪方法,以减少训练过程中的不确定性,从而使智能体能够快速适应新任务。
In the field of reinforcement learning (RL), agents are often tasked with solving a variety of problems differing only in their reward functions. In order to quickly obtain solutions to unseen problems with new reward functions, a popular approach involves functional composition of previously solved tasks. However, previous work using such functional composition has primarily focused on specific instances of composition functions whose limiting assumptions allow for exact zero-shot composition. Our work unifies these examples and provides a more general framework for compositionality in both standard and entropy-regularized RL. We find that, for a broad class of functions, the optimal solution for the composite task of interest can be related to the known primitive task solutions. Specifically, we present double-sided inequalities relating the optimal composite value function to the value functions for the primitive tasks. We also show that the regret of using a zero-shot policy can be bounded for this class of functions. The derived bounds can be used to develop clipping approaches for reducing uncertainty during training, allowing agents to quickly adapt to new tasks.