课题基金 / 基金详情

Curse-of-dimensionality-free nonlinear optimal feedback control with deep neural networks. A compositionality-based approach via Hamilton-Jacobi-Bellman PDEs

Curse-of-dimensionality-free nonlinear optimal feedback control with deep neural networks. A compositionality-based approach via Hamilton-Jacobi-Bellman PDEs
深度神经网络的无维数非线性最优反馈控制。
批准号:
463912816
负责人:
Professor Dr. Lars Grüne
金额:
$0.0万
依托单位国家:
德国
项目类别:
Priority Programmes
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:

项目摘要

项目成果

Professor Dr. Lars Grüne的其他基金

相似基金

相关文献

中文摘要
翻译
最优反馈控制是深度学习方法影响最大的领域之一。深度强化学习是获得最优反馈规律的方法之一,也可以说是人工智能中最成功的算法之一,它是人工智能在国际象棋或围棋等游戏中惊人表现的背后,也在科学、技术和经济中有着广泛的应用。从数学上讲,该方法背后的核心问题是如何通过深度神经网络(DNN)最好地表示最优值函数,即为每个状态分配最优性能值的函数,在强化学习中也称为成本-去函数。然后,可以从这些函数计算出最优反馈律。在连续时间内,这些最优值函数由Hamilton-Jacobi-Bellman偏微分方程组(HJB PDEs)表征,该方程通过DNN将问题与偏微分方程组的解联系起来。由于HJB偏微分方程组的维度由控制最优控制问题的动力学状态的维度决定,所以HJB方程自然形成了一类高维偏微分方程组。因此,它们倾向于众所周知的维度诅咒,即其解的数值努力在维度中呈指数增长的事实。众所周知,具有某些有益结构的函数,如组合函数或可分离函数,可以用具有适当结构的DNN来逼近,从而避免了维度灾难。对于刻画Lyapunov函数的HJB偏微分方程解,本项目的提出者最近证明了小增益条件--即关于问题动力学的特殊条件--建立了可分下解的存在性,DNN可以利用这些可分下解通过具有合适损失函数的训练算法来有效地逼近它们。这些结果为基于DNN的一般非线性HJB方程的无维数灾方法铺平了道路,这也是本项目的目标。除了小增益理论,还有一个大型工具箱的非线性反馈控制设计技术,导致合成(次)最优值函数。一方面,这些方法在数学上是可靠的,适用于许多现实世界的问题,但另一方面,当结果值函数或反馈规律需要计算时,它们带来了巨大的计算挑战。在这个项目中,我们将利用这些方法提供的结构洞察力来建立组成最优值函数或其近似的存在性,但通过使用适当的DNN训练算法来规避它们的计算复杂性。在这个过程中,我们将刻画最优反馈控制问题的特征,对于这些问题,通过DNN可以获得无维度诅咒(近似)解,并为计算这些解提供有效的网络结构和训练方案。
英文摘要
Optimal feedback control is one of the areas in which methods from deep learning have an enormous impact. Deep Reinforcement Learning, one of the methods for obtaining optimal feedback laws and arguably one of the most successful algorithms in artificial intelligence, stands behind the spectacular performance of artificial intelligence in games such as Chess or Go, but has also manifold applications in science, technology and economy. Mathematically, the core question behind this method is how to best represent optimal value functions, i.e., the functions that assign the optimal performance value to each state, also known as cost-to-go function in reinforcement learning, via deep neural networks (DNNs). The optimal feedback law can then be computed from these functions. In continuous time, these optimal value functions are characterised by Hamilton-Jacobi-Bellman partial differential equation (HJB PDEs), which links the question to the solution of PDEs via DNNs. As the dimension of the HJB PDE is determined by the dimension of the state of the dynamics governing the optimal control problem, HJB equations naturally form a class of high-dimensional PDEs. They are thus prone to the well-known curse of dimensionality, i.e., to the fact that the numerical effort for its solution grows exponentially in the dimension. It is known that functions with certain beneficial structures, like compositional or separable functions, can be approximated by DNNs with suitable architecture avoiding the curse of dimensionality. For HJB PDEs characterising Lyapunov functions it was recently shown by the proposer of this project that small-gain conditions - i.e., particular conditions on the dynamics of the problem - establish the existence of separable subsolutions, which can be exploited for efficiently approximating them by DNNs via training algorithms with suitable loss functions. These results pave the way for curse-of-dimensionality free DNN-based approaches for general nonlinear HJB equations, which are the goal of this project. Besides small-gain theory, there exists a large toolbox of nonlinear feedback control design techniques that lead to compositional (sub)optimal value functions. On the one hand, these methods are mathematically sound and apply to many real-world problems, but on the other hand they come with significant computational challenges when the resulting value functions or feedback laws shall be computed. In this project, we will exploit the structural insight provided these methods for establishing the existence of compositional optimal value functions or approximations thereof, but circumvent their computational complexity by using appropriate training algorithms for DNNs instead. Proceeding this way, we will characterise optimal feedback control problems for which curse-of-dimensionality-free (approximate) solutions via DNNs are possible and provide efficient network architectures and training schemes for computing these solutions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Specialized Adaptive Algorithms for Model Predictive Control of PDEs
  • 批准号:
    337928467
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2017
  • 负责人:
    Professor Dr. Lars Grüne
  • 依托单位:
Model predictive PDE control for energy efficient building operation:Economic model predictive control and time varying systems
  • 批准号:
    274853298
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2015
  • 负责人:
    Professor Dr. Lars Grüne
  • 依托单位:
Model Predictive Control for the Fokker-Planck Equation
Performance Analysis for Distributed and Multiobjective Model Predictive Control — The role of Pareto fronts, multiobjective dissipativity and multiple equilibria
  • 批准号:
    244602989
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2013
  • 负责人:
    Professor Dr. Lars Grüne
  • 依托单位:
国内基金
海外基金
高维稀疏数据聚类研究
  • 批准号:
    70771007
  • 项目类别:
    面上项目
  • 资助金额:
    16.0万元
  • 批准年份:
    2007
  • 负责人:
    武森
  • 依托单位: