Augmenting sampling based controllers with machine learning

Augmenting sampling based controllers with machine learning
复制标题

通过机器学习增强基于采样的控制器

DOI:
10.1145/3099564.3099579
复制
发表时间:
2017
期刊:
Proceedings of the ACM SIGGRAPH / Eurographics Symposium on Computer Animation
影响因子:
--
通讯作者:
Perttu Hämäläinen
Perttu Hämäläinen
中科院分区:
--
文献类型:
--
作者:
Joose Rajamäki;Perttu Hämäläinen

文献摘要

被引文献

相似文献

尽管最近在该领域取得了显著进展,但3D角色控制的有效学习仍然是一个悬而未决的问题。我们提出了一种将基于采样的模型预测控制器的规划与从计划控制中学习相结合的新算法。我们结合了两种学习方法:1)即时但不精确的最近邻学习,2)较慢但更精确的神经网络学习。最近邻学习允许快速锁定新的经验,而神经网络更渐进地学习并发展出稳定的数据表示。我们的实验表明,学习者相互增强,并允许快速发现和完善复杂的技能,如3D双足运动。我们在1、2和4腿的3D人物在干扰下的运动中证明了这一点,如沉重的弹丸击中和突然改变目标方向。当与学习器增强时,基于采样的模型预测控制器可以在4核CPU上在一分钟内产生这些稳定的步态。在训练期间,系统运行实时或交互帧率取决于字符的复杂性。
Efficient learning of 3D character control still remains an open problem despite of the remarkable recent advances in the field. We propose a new algorithm that combines planning by a sampling-based model-predictive controller and learning from the planned control, which is very noisy. We combine two methods of learning: 1) immediate but imprecise nearest-neighbor learning, and 2) slower but more precise neural net-work learning. The nearest neighbor learning allows to rapidly latch on to new experiences whilst the neural net-work learns more gradually and develops a stable representation of the data. Our experiments indicate that the learners augment each other, and allow rapid discovery and refinement of complex skills such as 3D bipedal locomotion. We demonstrate this in locomotion of 1-, 2- and 4-legged 3D characters under disturbances such as heavy projectile hits and abruptly changing target direction. When augmented with the learners, the sampling based model predictive controller can produce these stable gaits in under a minute on a 4-core CPU. During training the system runs real-time or at interactive frame rates depending on the character complexity.