Sample-Efficient Multimodal Dynamics Modeling for Risk-Sensitive Reinforcement Learning
Sample-Efficient Multimodal Dynamics Modeling for Risk-Sensitive Reinforcement Learning
复制标题
用于风险敏感强化学习的样本高效多模态动力学建模
DOI:
10.1109/icmre54455.2022.9734091
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Hashimoto Koichi
中科院分区:
文献类型:
--
作者:
Yashima Ryota;Yamaguchi Akihiko;Hashimoto Koichi
When we consider dynamics of pouring viscous liquid, there are multiple modes; e.g., liquid jams when the size of container opening is narrow. Our past work showed that using multiple skills (e.g., tipping and shaking a container) is effective to handle such tasks where stochastic neural networks were used to model the multimodal dynamics in a model-based control. It was possible because we could assume the output-state probability distribution is unimodal for an input state-action distribution. However, we have found that the output-state distribution becomes multimodal where the mode of dynamics switches, and the prediction of the stochastic neural networks becomes inaccurate due to the unimodal assumption. As the consequence, the model-based control may choose a risky skill and parameters. This paper explores modeling such multimodal dynamics. Since the output distribution becomes multimodal only where the dynamics mode switches, the number of available samples might be limited. Furthermore, we do not need to explicitly handle the output multimodality since our goal is a stochastic model-based control (dynamic programming). Thus, we propose to introduce a Gaussian mixture model to expand the variance of output distribution of the stochastic neural networks. This model can be easily unified into existing stochastic dynamic programming. The simulation experiments of pouring viscous liquid demonstrated that the proposed method improves the risk sensitivity.