Sparse Gaussian Process Temporal Difference Learning for Marine Robot Navigation

Sparse Gaussian Process Temporal Difference Learning for Marine Robot Navigation
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
--
影响因子:
--
通讯作者:
John D. Martin;Jinkun Wang;Brendan Englot
John D. Martin;Jinkun Wang;Brendan Englot
中科院分区:
其他
文献类型:
--
作者:
John D. Martin;Jinkun Wang;Brendan Englot

文献摘要

相似文献

我们提出了一种时间差(TD)学习的方法,解决了机器人学习在海洋环境中导航所面临的几个挑战。为了提高数据效率,我们的方法减少了对高斯过程回归的TD更新。为了使预测服从在线设置,我们引入了一个稀疏近似,其质量优于当前基于拒绝的稀疏方法。我们推导出预测值函数的后验和使用的时刻,以获得一个新的算法,无模型的政策评估,SPGP-SARSA。通过简单的修改,我们表明SPGP-SARSA可以简化为基于模型的等价物SPGP-TD。我们进行全面的模拟研究,并与水下机器人进行物理学习试验。我们的研究结果表明,SPGP-SARSA可以优于最先进的稀疏方法,复制其精确对应的预测质量,并适用于解决水下导航任务。
We present a method for Temporal Difference (TD) learning that addresses several challenges faced by robots learning to navigate in a marine environment. For improved data efficiency, our method reduces TD updates to Gaussian Process regression. To make predictions amenable to online settings, we introduce a sparse approximation with improved quality over current rejection-based sparse methods. We derive the predictive value function posterior and use the moments to obtain a new algorithm for model-free policy evaluation, SPGP-SARSA. With simple changes, we show SPGP-SARSA can be reduced to a model-based equivalent, SPGP-TD. We perform comprehensive simulation studies and also conduct physical learning trials with an underwater robot. Our results show SPGP-SARSA can outperform the state-of-the-art sparse method, replicate the prediction quality of its exact counterpart, and be applied to solve underwater navigation tasks.