Approximating gradients for differentiable quality diversity in reinforcement learning

Approximating gradients for differentiable quality diversity in reinforcement learning
复制标题

DOI:
10.1145/3512290.3528705
复制
发表时间:
2022-02
期刊:
Proceedings of the Genetic and Evolutionary Computation Conference
影响因子:
--
通讯作者:
Bryon Tjanaka;Matthew C. Fontaine;J. Togelius;S. Nikolaidis
Bryon Tjanaka;Matthew C. Fontaine;J. Togelius;S. Nikolaidis
中科院分区:
其他
文献类型:
--
作者:
Bryon Tjanaka;Matthew C. Fontaine;J. Togelius;S. Nikolaidis

文献摘要

相似文献

考虑训练健壮的智能体的问题。一种方法是生成不同的代理策略集合。然后可以将训练视为质量多样性(QD)优化问题,我们在其中搜索与量化行为相关的多种性能策略的集合。最近的研究表明,当精确梯度可用时,可微分质量分集(DQD)算法大大加快了QD优化。然而,代理策略通常假定环境不可微分。为了将DQD算法应用于训练代理策略,我们必须近似性能和行为的梯度。我们提出了当前最先进的DQD算法的两个变体,它们通过强化学习(RL)中常见的近似方法计算梯度。我们在四个模拟运动任务中评估了我们的方法。一种变体在结合QD和RL方面取得了与当前最先进的成果相当的结果,而另一种变体在两种运动任务中表现相当。这些结果提供了深入了解当前DQD算法在梯度必须近似的领域的局限性。源代码可从https://github.com/icaros-usc/dqd-rl获得
Consider the problem of training robustly capable agents. One approach is to generate a diverse collection of agent polices. Training can then be viewed as a quality diversity (QD) optimization problem, where we search for a collection of performant policies that are diverse with respect to quantified behavior. Recent work shows that differentiable quality diversity (DQD) algorithms greatly accelerate QD optimization when exact gradients are available. However, agent policies typically assume that the environment is not differentiable. To apply DQD algorithms to training agent policies, we must approximate gradients for performance and behavior. We propose two variants of the current state-of-the-art DQD algorithm that compute gradients via approximation methods common in reinforcement learning (RL). We evaluate our approach on four simulated locomotion tasks. One variant achieves results comparable to the current state-of-the-art in combining QD and RL, while the other performs comparably in two locomotion tasks. These results provide insight into the limitations of current DQD algorithms in domains where gradients must be approximated. Source code is available at https://github.com/icaros-usc/dqd-rl