How Are Learned Perception-Based Controllers Impacted by the Limits of Robust Control?

How Are Learned Perception-Based Controllers Impacted by the Limits of Robust Control?
复制标题

DOI:
--
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Jingxi Xu;Bruce Lee;N. Matni;Dinesh Jayaraman
Jingxi Xu;Bruce Lee;N. Matni;Dinesh Jayaraman
中科院分区:
其他
文献类型:
--
作者:
Jingxi Xu;Bruce Lee;N. Matni;Dinesh Jayaraman

文献摘要

被引文献

相似文献

最优控制问题的难度通常用系统性质来描述,例如能控性/能观测性的最小特征值。在强化学习(RL)等数据驱动技术日益流行的背景下,以及在输入观测是高维图像且转换动力学未知的控制环境中,我们重新审视了这些特征。具体地说,我们问:任务的可量化控制和感知难度指标在多大程度上预测了数据驱动控制器的性能和样本复杂性?我们将两种不同类型的局部可观测性调制到手杖“棍棒平衡”问题中--(I)手杖上一个可见固定点的高度,它可用于调整任何控制器可实现的性能的基本极限,以及(Ii)从手杖的深度或RGB图像推断出的在固定点位置的感知噪声水平。在这些背景下,我们使用视觉估计的系统状态对两类流行的控制器:RL和基于系统识别的$H\INFTY$CONTROL进行实证研究。我们的结果表明,鲁棒控制的基本极限对基于感知的学习控制器的采样效率和性能有相应的影响。有关更多信息,请访问我们的项目网站https://jxu.ai/rl-vs-control-web。
The difficulty of optimal control problems has classically been characterized in terms of system properties such as minimum eigenvalues of controllability/observability gramians. We revisit these characterizations in the context of the increasing popularity of data-driven techniques like reinforcement learning (RL), and in control settings where input observations are high-dimensional images and transition dynamics are unknown. Specifically, we ask: to what extent are quantifiable control and perceptual difficulty metrics of a task predictive of the performance and sample complexity of data-driven controllers? We modulate two different types of partial observability in a cartpole"stick-balancing"problem -- (i) the height of one visible fixation point on the cartpole, which can be used to tune fundamental limits of performance achievable by any controller, and by (ii) the level of perception noise in the fixation point position inferred from depth or RGB images of the cartpole. In these settings, we empirically study two popular families of controllers: RL and system identification-based $H_\infty$ control, using visually estimated system state. Our results show that the fundamental limits of robust control have corresponding implications for the sample-efficiency and performance of learned perception-based controllers. Visit our project website https://jxu.ai/rl-vs-control-web for more information.