Fundamental Performance Limits for Sensor-Based Robot Control and Policy Learning

Fundamental Performance Limits for Sensor-Based Robot Control and Policy Learning
复制标题

DOI:
10.15607/rss.2022.xviii.036
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Anirudha Majumdar;Vincent Pacelli
Anirudha Majumdar;Vincent Pacelli
中科院分区:
其他
文献类型:
--
作者:
Anirudha Majumdar;Vincent Pacelli

文献摘要

相似文献

-我们的目标是开发理论和算法,以建立机器人传感器对给定任务的性能的基本限制。为了实现这一点,我们得fiNe一个量,它捕获由传感器提供的与任务相关的信息量。利用信息论中的广义Fano不等式的一个新形式,我们证明了这个量提供了一步决策任务最高可达期望报酬的一个上界。然后,我们通过动态规划方法将这个界扩展到多步问题。我们给出了数值计算结果界的算法,并在三个例子上演示了我们的方法:(I)文献中关于部分可观测的马尔可夫决策过程的LAVA问题,(Ii)与机器人捕捉自由落体物体相对应的具有连续状态和观测空间的例子,以及(Iii)使用具有非高斯噪声的深度传感器的避障。我们通过比较我们的上界和可达到的下界(通过合成或学习具体的控制策略来计算)来证明我们的方法能够为这些问题建立可实现的性能的强极限。
—Our goal is to develop theory and algorithms for establishing fundamental limits on performance for a given task imposed by a robot’s sensors. In order to achieve this, we define a quantity that captures the amount of task-relevant information provided by a sensor. Using a novel version of the generalized Fano inequality from information theory, we demonstrate that this quantity provides an upper bound on the highest achievable expected reward for one-step decision making tasks. We then extend this bound to multi-step problems via a dynamic programming approach. We present algorithms for numerically computing the resulting bounds, and demonstrate our approach on three examples: (i) the lava problem from the literature on partially observable Markov decision processes, (ii) an example with continuous state and observation spaces corresponding to a robot catching a freely-falling object, and (iii) obstacle avoidance using a depth sensor with non-Gaussian noise. We demonstrate the ability of our approach to establish strong limits on achievable performance for these problems by comparing our upper bounds with achievable lower bounds (computed by synthesizing or learning concrete control policies).