PFPN: Continuous Control of Physically Simulated Characters using Particle Filtering Policy Network

PFPN: Continuous Control of Physically Simulated Characters using Particle Filtering Policy Network
复制标题

DOI:
10.1145/3487983.3488301
复制
发表时间:
2020-03
期刊:
Proceedings of the 14th ACM SIGGRAPH Conference on Motion, Interaction and Games
影响因子:
--
通讯作者:
Pei Xu;Ioannis Karamouzas
Pei Xu;Ioannis Karamouzas
中科院分区:
其他
文献类型:
--
作者:
Pei Xu;Ioannis Karamouzas

文献摘要

相似文献

使用强化学习的基于物理的角色控制的数据驱动方法已经成功地应用于生成高质量的运动。然而,现有的方法通常依赖于高斯分布来表示动作策略,这在解决高关节角色的高维连续控制问题时可能过早地承诺次优动作。为了提高基于物理的角色控制器的学习性能,我们提出了一种基于粒子的动作策略来代替高斯策略的框架。我们利用粒子滤波对动作空间进行动态探索和离散化,并跟踪表现为混合分布的后验策略。所得到的策略可以取代角色控制问题的主要单峰高斯策略,而不改变用于执行策略优化的强化学习算法的底层模型体系结构。我们展示了我们的方法在各种运动捕捉模拟任务中的适用性。与使用高斯的相应实现相比,使用我们的基于粒子的策略的基线获得了更好的模拟性能和收敛速度,并且在角色控制期间对外部扰动更稳健。相关代码可在以下网址获得:https://motion-lab.github.io/PFPN.
Data-driven methods for physics-based character control using reinforcement learning have been successfully applied to generate high-quality motions. However, existing approaches typically rely on Gaussian distributions to represent the action policy, which can prematurely commit to suboptimal actions when solving high-dimensional continuous control problems for highly-articulated characters. In this paper, to improve the learning performance of physics-based character controllers, we propose a framework that considers a particle-based action policy as a substitute for Gaussian policies. We exploit particle filtering to dynamically explore and discretize the action space, and track the posterior policy represented as a mixture distribution. The resulting policy can replace the unimodal Gaussian policy which has been the staple for character control problems, without changing the underlying model architecture of the reinforcement learning algorithm used to perform policy optimization. We demonstrate the applicability of our approach on various motion capture imitation tasks. Baselines using our particle-based policies achieve better imitation performance and speed of convergence as compared to corresponding implementations using Gaussians, and are more robust to external perturbations during character control. Related code is available at: https://motion-lab.github.io/PFPN.