Estimation and Control Using Sampling-Based Bayesian Reinforcement Learning

Estimation and Control Using Sampling-Based Bayesian Reinforcement Learning
复制标题

DOI:
10.1049/iet-cps.2019.0045
复制
发表时间:
2018-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Patrick Slade;Preston Culbertson;Zachary Sunberg;Mykel J. Kochenderfer
Patrick Slade;Preston Culbertson;Zachary Sunberg;Mykel J. Kochenderfer
中科院分区:
其他
文献类型:
--
作者:
Patrick Slade;Preston Culbertson;Zachary Sunberg;Mykel J. Kochenderfer

文献摘要

被引文献

相似文献

现实世界的自主系统在其姿态和动力学的不确定性下运行。自主控制系统必须同时执行估计和控制任务,以保持对动态变化或建模误差的鲁棒性。然而,信息收集行动往往与达到控制目标的最佳行动相冲突,需要在探索和开发之间进行权衡。这里考虑的具体问题设置是离散时间非线性系统,过程噪声,输入约束,参数不确定性。本文将此问题框架为贝叶斯自适应马尔可夫决策过程,并使用蒙特卡洛树搜索和无迹卡尔曼滤波器在线解决该问题,以考虑过程噪声和参数不确定性。该方法与确定性等效模型预测控制和近似QMDP解决方案的树搜索方法进行了比较,提供了信息收集有用时的洞察。离散时间仿真表征了在一系列过程噪声和未知参数的范围内的性能。一个离线优化方法被用来选择蒙特卡洛树搜索参数,而无需手动调整。代替递归的可行性保证,提供了一个概率边界启发式,增加了保持在所需区域内的状态的概率。
Real-world autonomous systems operate under uncertainty about both their pose and dynamics. Autonomous control systems must simultaneously perform estimation and control tasks to maintain robustness to changing dynamics or modeling errors. However, information gathering actions often conflict with optimal actions for reaching control objectives, requiring a trade-off between exploration and exploitation. The specific problem setting considered here is for discrete-time nonlinear systems, with process noise, input-constraints, and parameter uncertainty. This article frames this problem as a Bayes-adaptive Markov decision process and solves it online using Monte Carlo tree search with an unscented Kalman filter to account for process noise and parameter uncertainty. This method is compared with certainty equivalent model predictive control and a tree search method that approximates the QMDP solution, providing insight into when information gathering is useful. Discrete time simulations characterize performance over a range of process noise and bounds on unknown parameters. An offline optimization method is used to select the Monte Carlo tree search parameters without hand-tuning. In lieu of recursive feasibility guarantees, a probabilistic bounding heuristic is offered that increases the probability of keeping the state within a desired region.