Learning control lyapunov functions from counterexamples and demonstrations

Learning control lyapunov functions from counterexamples and demonstrations
复制标题

DOI:
10.1007/s10514-018-9791-9
复制
发表时间:
2019-02-01
期刊:
影响因子:
3.5
通讯作者:
Sankaranarayanan, Sriram
Sankaranarayanan, Sriram
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ravanbakhsh, Hadi;Sankaranarayanan, Sriram

文献摘要

被引文献

相似文献

我们提出了一种学习控制Lyapunov的技术,该技术依次用于合成可以稳定系统的非线性动力学系统的控制器,或满足诸如安全集合内部的规格,或者最终达到目标集,同时保持在内部。安全套装。学习框架使用了一个演示者,该示威者实现了一个黑框,不受信任的策略来解决感兴趣的问题,这是一个对示威者有限的查询来推断候选功能的学习者,以及检查当前候选人是否是一个验证者有效的控制Lyapunov样函数。总体学习框架是迭代的,使用验证者发现的反例和这些反例上的示范消除了每次迭代的一组候选者。我们使用凸优化的椭圆形近似技术证明了它的收敛性。我们还使用非线性MPC控制器实施了此方案,以作为非线性动力学系统的一组状态和轨迹稳定问题的演示者。我们展示了如何使用多项式系统的验证问题的凸松弛来有效地构建验证者,以进行半准编程问题实例。我们的方法能够综合相对简单的多项式控制类似Lyapunov的功能,在此过程中,使用保证和计算较便宜的控制器替换MPC。
We present a technique for learning control Lyapunov-like functions, which are used in turn to synthesize controllers for nonlinear dynamical systems that can stabilize the system, or satisfy specifications such as remaining inside a safe set, or eventually reaching a target set while remaining inside a safe set. The learning framework uses a demonstrator that implements a black-box, untrusted strategy presumed to solve the problem of interest, a learner that poses finitely many queries to the demonstrator to infer a candidate function, and a verifier that checks whether the current candidate is a valid control Lyapunov-like function. The overall learning framework is iterative, eliminating a set of candidates on each iteration using the counterexamples discovered by the verifier and the demonstrations over these counterexamples. We prove its convergence using ellipsoidal approximation techniques from convex optimization. We also implement this scheme using nonlinear MPC controllers to serve as demonstrators for a set of state and trajectory stabilization problems for nonlinear dynamical systems. We show how the verifier can be constructed efficiently using convex relaxations of the verification problem for polynomial systems to semi-definite programming problem instances. Our approach is able to synthesize relatively simple polynomial control Lyapunov-like functions, and in that process replace the MPC using a guaranteed and computationally less expensive controller.