课题基金 / 基金详情

Online Learning-based Real-time Control of Unknown Autonomous Systems

Online Learning-based Real-time Control of Unknown Autonomous Systems
基于在线学习的未知自治系统实时控制
批准号:
1810447
负责人:
Rahul Jain
金额:
$33.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-15 至 2022-07-31

项目摘要

项目成果

Rahul Jain的其他基金

相似基金

相关文献

中文摘要
翻译
许多新兴的自治系统,例如,非结构化环境中的机器人太复杂,无法精确建模。存在未知的模型参数、部分状态观测值或系统特性的漂移。这使得系统识别和控制的问题相当具有挑战性。需要实时适应以实现最佳和弹性操作。众所周知,经典的自适应控制方法的系统辨识和“确定性等效”控制的反馈回路是行不通的。在这个项目中,我们介绍了一个新的范例的“学习控制”未知自治系统的基础上,新发展的汤普森/后验采样为基础的在线学习方法。我们将重点讨论马尔可夫决策过程(MDP)的离散状态空间模型。首先,我们将开发一个后采样启发算法的在线学习为基础的控制与实时适应MDP模型与部分观察系统状态。我们注意到,这种方法可以解释为提供恰到好处的随机化,以最佳地权衡探索和开发,这是以最快的速度在线学习最优策略所需的。然后,我们将扩展到系统参数可能随时间变化或漂移的设置。然后,我们将开发这样的算法更相关,但也更复杂的系统模型-随机混合系统,既有离散和连续状态。开发的算法将在OpenAI Gym的经典控制和机器人环境中的模拟实验中得到广泛验证。该研究的智力价值在于其对“自治系统科学”的贡献,它通过解决关于分离各种随机系统模型的参数估计,状态估计和控制的基本问题,特别是当模型参数必须从数据中学习时,为自治系统开发基于在线学习的实时控制和自适应基础。更广泛的影响将包括通过传播研究成果、培训一名女博士生和K-12 STEM外展工作对智能电网、自主机器人和医疗CPS设备的影响。该奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Many emerging autonomous systems, e.g., robots in unstructured environments, are too complex to be accurately modeled. There are unknown model parameters, partial state observations, or a drift in system characteristics. This makes the problem of system identification and control quite challenging. Real-time adaptation is needed for optimal and resilient operation. It is well-known that the classical adaptive con-trol approach of system identification and `certainty equivalent' control in the feedback-loop doesn't work. In this project, we introduce a new paradigm of 'Learning-to-Control' unknown Autonomous Systems based on the newly developing approach of Thompson/Posterior sampling-based online learning. We will focus on discrete state space models of Markov decision processes (MDPs). We will first develop a posterior sampling-inspired algorithms for online learning-based control with real-time adaptation for MDP models with partial observation of the system state. We note that such approaches may be inter-preted to provide just the right amount of randomization for optimally trading off exploration and exploi-tation that is needed for online learning of the optimal policy at the fastest rate. We will then extend this to the setting where the system parameter may be varying or drifting with time. We will then develop such algorithms for more relevant but also more complicated system models - stochastic hybrid systems, that have both discrete and continuous states. The developed algorithms will be extensively validated in sim-ulation experiments in the classical control and robotics environments in OpenAI Gym. The intellectual merit of the research lies in its contribution to the 'Science of Autonomous Systems' by development of foundations of online learning-based real-time control and adaptation for autonomous systems by addressing fundamental questions about separation of parameter estimation, state estima-tion and control for various stochastic system models, particularly when model parameters must be learnt from data. The broader impacts will include impact on the smart grid, autonomous robotics, and medical CPS devices via dissemination of research results, training of a female PhD student and a K-12 STEM outreach effort.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(11)
专著(0)
科研奖励(0)
会议论文
Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes”, ICML (Int'l Conf. on Machine Learning) 2020.
无限视野平均奖励马尔可夫决策过程中的无模型强化学习,ICML(机器学习国际会议)2020。
DOI: --
发表时间: 2020
期刊: Proceedings of Machine Learning Research
影响因子: --
作者: [Chen-Yu Wei, Mehdi Jafarnia-Jahromi]
通讯作者: Chen-Yu Wei, Mehdi Jafarnia-Jahromi
DOI: 10.1609/aaai.v35i9.16979
发表时间: 2020-09
期刊: ArXiv
影响因子: --
作者: [K. C. Kalagarla;Rahul Jain;P. Nuzzo]
通讯作者: K. C. Kalagarla;Rahul Jain;P. Nuzzo
Scheduling Flexible Nonpreemptive Loads in Smart-Grid Networks
智能电网网络中灵活的非抢占式负载调度
DOI: 10.1109/tcns.2022.3141017
发表时间: 2022
期刊: IEEE Transactions on Control of Network Systems
影响因子: 4.2
作者: [Dahlin, Nathan, Jain, Rahul]
通讯作者: Jain, Rahul
DOI: 10.1016/j.automatica.2020.109016
发表时间: 2017-08
期刊: Autom.
影响因子: --
作者: [Mehdi Jafarnia-Jahromi;Rahul Jain]
通讯作者: Mehdi Jafarnia-Jahromi;Rahul Jain
11
    EAGER: Real-Time: Formal Reinforcement Learning Methods for the Design of Safety-critical Autonomous Systems
    • 批准号:
      1839842
    • 项目类别:
      Standard Grant
    • 资助金额:
      $28.61万
    • 财政年份:
      2019
    • 负责人:
      Rahul Jain
    • 依托单位:
    AF: Small: A New Approach to Analysis and Design of Algorithms for Stochastic Control and Optimization
    • 批准号:
      1817212
    • 项目类别:
      Standard Grant
    • 资助金额:
      $40.0万
    • 财政年份:
      2018
    • 负责人:
      Rahul Jain
    • 依托单位:
    Collaborative Research: Smarter Markets for a Smarter Grid: Pricing Randomness, Flexibility and Risk
    • 批准号:
      1611574
    • 项目类别:
      Standard Grant
    • 资助金额:
      $22.5万
    • 财政年份:
      2016
    • 负责人:
      Rahul Jain
    • 依托单位:
    CAREER: Network Economics: Theory and Architectures for Incentive-engineered Networks
    • 批准号:
      0954116
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $42.5万
    • 财政年份:
      2010
    • 负责人:
      Rahul Jain
    • 依托单位:
    国内基金
    海外基金
    Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
    Understanding structural evolution of galaxies with machine learning
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      Nicola Rosario Napolitano
    • 依托单位:
    煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
    • 批准号:
      --
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      30万元
    • 批准年份:
      2022
    • 负责人:
      吉建娇
    • 依托单位:
    基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
    • 批准号:
      62003314
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      24.0万元
    • 批准年份:
      2020
    • 负责人:
      沈剑
    • 依托单位: