Sample-Efficient Multimodal Dynamics Modeling for Risk-Sensitive Reinforcement Learning

Sample-Efficient Multimodal Dynamics Modeling for Risk-Sensitive Reinforcement Learning
复制标题

用于风险敏感强化学习的样本高效多模态动力学建模

DOI:
10.1109/icmre54455.2022.9734091
复制
发表时间:
2022
期刊:
2022 8th International Conference on Mechatronics and Robotics Engineering (ICMRE)
影响因子:
--
通讯作者:
Hashimoto Koichi
Hashimoto Koichi
中科院分区:
--
文献类型:
--
作者:
Yashima Ryota;Yamaguchi Akihiko;Hashimoto Koichi

文献摘要

相似文献

当我们考虑倾倒粘性液体的动力学时,存在多种模式;例如,当容器开口的尺寸窄时,液体堵塞。我们过去的工作表明,使用多种技能(例如,倾斜和摇动容器)对于处理这样的任务是有效的,其中随机神经网络被用于在基于模型的控制中对多模态动态进行建模。这是可能的,因为我们可以假设输出-状态概率分布对于输入状态-动作分布是单峰的。然而,我们已经发现,输出状态分布成为多模态的动态模式切换,和随机神经网络的预测变得不准确,由于单峰假设。因此,基于模型的控制可能选择有风险的技能和参数。本文探讨了这种多模态动力学建模。由于输出分布仅在动态模式切换的情况下变为多模态,因此可用样本的数量可能有限。此外,我们不需要显式地处理输出多模态,因为我们的目标是基于随机模型的控制(动态规划)。因此,我们建议引入高斯混合模型来扩大随机神经网络输出分布的方差。该模型可以很容易地统一到现有的随机动态规划。粘性液体浇注仿真实验表明,该方法提高了风险敏感性。
When we consider dynamics of pouring viscous liquid, there are multiple modes; e.g., liquid jams when the size of container opening is narrow. Our past work showed that using multiple skills (e.g., tipping and shaking a container) is effective to handle such tasks where stochastic neural networks were used to model the multimodal dynamics in a model-based control. It was possible because we could assume the output-state probability distribution is unimodal for an input state-action distribution. However, we have found that the output-state distribution becomes multimodal where the mode of dynamics switches, and the prediction of the stochastic neural networks becomes inaccurate due to the unimodal assumption. As the consequence, the model-based control may choose a risky skill and parameters. This paper explores modeling such multimodal dynamics. Since the output distribution becomes multimodal only where the dynamics mode switches, the number of available samples might be limited. Furthermore, we do not need to explicitly handle the output multimodality since our goal is a stochastic model-based control (dynamic programming). Thus, we propose to introduce a Gaussian mixture model to expand the variance of output distribution of the stochastic neural networks. This model can be easily unified into existing stochastic dynamic programming. The simulation experiments of pouring viscous liquid demonstrated that the proposed method improves the risk sensitivity.