Variational Inference MPC using Tsallis Divergence

Variational Inference MPC using Tsallis Divergence
复制标题

DOI:
10.15607/rss.2021.xvii.073
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Ziyi Wang;Oswin So;Jason Gibson;Bogdan I. Vlahov;Manan S. Gandhi;Guan-Horng Liu;Evangelos A. Theodorou
Ziyi Wang;Oswin So;Jason Gibson;Bogdan I. Vlahov;Manan S. Gandhi;Guan-Horng Liu;Evangelos A. Theodorou
中科院分区:
其他
文献类型:
--
作者:
Ziyi Wang;Oswin So;Jason Gibson;Bogdan I. Vlahov;Manan S. Gandhi;Guan-Horng Liu;Evangelos A. Theodorou

文献摘要

相似文献

在本文中,我们通过使用非扩展 Tsallis 散度为变分推理-随机最优控制提供了一个通用框架。通过将变形指数函数融入最优似然函数,推导了一种新颖的Tsallis变分推理模型预测控制算法,其中包括变分推理模型预测控制、模型预测路径积分控制、交叉熵方法和Stein变分推理模型预测控制等先前工作作为特例。所提出的算法允许有效控制成本/回报变换,并且在相关成本的均值和方差减少方面具有优异的性能。上述特征得到了对所提出算法的风险敏感性水平的理论和数值分析以及对具有 3 种不同策略参数化的 5 个不同机器人系统的模拟实验的支持。
In this paper, we provide a generalized framework for Variational Inference-Stochastic Optimal Control by using thenon-extensive Tsallis divergence. By incorporating the deformed exponential function into the optimality likelihood function, a novel Tsallis Variational Inference-Model Predictive Control algorithm is derived, which includes prior works such as Variational Inference-Model Predictive Control, Model Predictive PathIntegral Control, Cross Entropy Method, and Stein VariationalInference Model Predictive Control as special cases. The proposed algorithm allows for effective control of the cost/reward transform and is characterized by superior performance in terms of mean and variance reduction of the associated cost. The aforementioned features are supported by a theoretical and numerical analysis on the level of risk sensitivity of the proposed algorithm as well as simulation experiments on 5 different robotic systems with 3 different policy parameterizations.