Adaptive Ensemble Q-learning: Minimizing Estimation Bias via Error Feedback

Adaptive Ensemble Q-learning: Minimizing Estimation Bias via Error Feedback
复制标题

DOI:
10.48550/arxiv.2306.11918
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Hang Wang;Sen Lin;Junshan Zhang
Hang Wang;Sen Lin;Junshan Zhang
中科院分区:
其他
文献类型:
--
作者:
Hang Wang;Sen Lin;Junshan Zhang

文献摘要

被引文献

相似文献

集成方法是一种很有前途的方法来减轻Q学习中的高估问题,其中使用多个函数逼近器来估计动作值。已知估计偏差严重依赖于系综大小(即,在目标中使用的Q函数逼近器的数量),并且由于在学习过程中函数逼近误差的时变性质,确定“正确的”系综大小是非常重要的。为了解决这一挑战,我们首先推导出估计偏差的上界和下界,基于该上界和下界,系综大小被适配为驱动偏差接近于零,从而相应地应对时变近似误差的影响。受理论研究结果的启发,我们主张将集成方法与模型辨识自适应控制(MIAC)相结合,以实现有效的集成规模自适应。具体而言,我们设计自适应Ensemble Q学习(AdaEQ),广义集成方法有两个关键步骤:(a)近似误差表征,作为灵活控制集成大小的反馈,和(B)集成大小自适应最小化的估计偏差。大量的实验表明,AdaEQ可以提高学习性能比现有的方法为MuJoCo基准。
The ensemble method is a promising way to mitigate the overestimation issue in Q-learning, where multiple function approximators are used to estimate the action values. It is known that the estimation bias hinges heavily on the ensemble size (i.e., the number of Q-function approximators used in the target), and that determining the `right' ensemble size is highly nontrivial, because of the time-varying nature of the function approximation errors during the learning process. To tackle this challenge, we first derive an upper bound and a lower bound on the estimation bias, based on which the ensemble size is adapted to drive the bias to be nearly zero, thereby coping with the impact of the time-varying approximation errors accordingly. Motivated by the theoretic findings, we advocate that the ensemble method can be combined with Model Identification Adaptive Control (MIAC) for effective ensemble size adaptation. Specifically, we devise Adaptive Ensemble Q-learning (AdaEQ), a generalized ensemble method with two key steps: (a) approximation error characterization which serves as the feedback for flexibly controlling the ensemble size, and (b) ensemble size adaptation tailored towards minimizing the estimation bias. Extensive experiments are carried out to show that AdaEQ can improve the learning performance than the existing methods for the MuJoCo benchmark.