Developing robust and scalable reinforcement learning algorithms
Developing robust and scalable reinforcement learning algorithms
批准号:
2740739
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
强化学习(RL)代表了一种将机器学习应用于复杂决策任务的强大范例。然而,当前的RL算法是脆弱的,并且有许多失败机制,使得它们很难应用到研究中使用的合成简单环境之外。在我的博士学位期间,我的目标是帮助开发健壮和可伸缩的RL算法,以便RL可以更容易地应用于大型、复杂、现实世界的问题。其中一个有前途的方法是离线RL;它是一套广泛的方法,利用预先存在的数据集来消除标准RL方法所需的在线数据收集需求。大规模数据集的使用一直是最近监督学习中变革性进展的驱动力,而离线RL提供了一种在RL中实现类似缩放过程的方法。除此之外,在许多在线数据收集成本高昂或存在安全风险的真实世界领域,离线RL是一个实用的选择,例如在机器人或医疗保健领域。然而,离线方法引入了额外的问题,即需要评估数据集中未涵盖的操作或行为的有效性。这一额外的挑战已经被证明是导致现有在线RL算法失败的原因,因此是健壮的、可扩展的离线RL算法的重要障碍。在我最近的项目中,我们研究了这个问题,并在基于离线模型的强化学习中确定了“到达边缘”的失败机制。基于这些见解,我们能够提出一种“价值不确定性感知”的方法来解决这个问题。作为扩展,我计划调查元学习是否可以用来发现一种更有效地处理值不确定性的离线RL算法。大规模RL的另一个障碍是“可塑性损失”现象,由此RL训练中固有的非平稳性已被证明导致深层RL代理逐渐失去从新数据中学习的能力。我目前正在调查这种现象背后的潜在机制,希望这将有助于开发更有效的缓解技术。最后,最近大型语言模型的进展为将语言作为工具向RL代理灌输现实世界的知识和归纳偏见提供了巨大的机会。这种整合的最有效方法是一个悬而未决的问题,我计划致力于探索这一点。总之,RL具有当前方法尚未实现的巨大潜力,在我的博士学位期间,我的目标是为开发RL算法做出贡献,这些算法是健壮的、可伸缩的,并且在应用于复杂的现实世界问题时是有效的。
英文摘要
Reinforcement learning (RL) represents a powerful paradigm for applying machine learning to complex decision making tasks. However, current RL algorithms are brittle and have numerous failure mechanisms, making them difficult to apply to beyond the synthetic simple environments used in research. Over my PhD I aim to help in developing robust and scalable RL algorithms, such that RL can be more easily applied to large, complex, real-world problems.One promising approach to this is offline RL; a broad set of methods which utilize pre-existing datasets to remove the need for the online data collection required by standard RL methods. Use of large-scale datasets has been the driver of the transformative recent advances in supervised learning, and offline RL presents a way of enabling similar scaling progress in RL. In addition to this, offline RL is a practical choice in many real-world domains where online data collection is costly or poses safety risks, such as in robotics or healthcare. The offline approach, however, introduces the extra problem of needing to evaluate effectiveness of actions or behaviour not covered in the dataset. This additional challenge has been shown to cause existing online RL algorithms to fail, and hence represents a significant barrier to robust, scalable offline RL algorithms.In my recent project we investigated this issue and identified the "edge-of-reach" failure mechanism in offline model-based reinforcement learning. Based on these insights we were able propose a "value uncertainty-aware" approach which resolves this issue. As an extension to this, I plan to investigate whether meta-learning can be used to discover an offline RL algorithm that deals with value uncertainty even more effectively.Another barrier to large-scale RL is the phenomenon of "plasticity loss," whereby the inherent non-stationarity in RL training has been shown to cause deep RL agents to gradually lose the ability to learn from new data. I am currently investigating the underlying mechanisms behind this phenomenon, with the hope that this will inform development of more effective mitigation techniques.Finally, the recent advances in large language models present huge opportunity for using language as a tool for instilling real-world knowledge and inductive biases into RL agents. The most effective approach for this integration is an open question, and I plan to work on exploring this. In summary, RL has huge potential which is not yet realized by current approaches, and during my PhD I aim to contribute towards developing RL algorithms which are robust and scalable and are effective when applied to complex real-world problems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
半定松弛与非凸二次约束二次规划研究
-
批准号:11271243
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2012
-
负责人:王燕军
-
依托单位:
基于复合编码脉冲串的水下主动隐蔽性探测新方法研究
-
批准号:61271414
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2012
-
负责人:冯西安
-
依托单位:
民航客运网络收益管理若干问题的研究
-
批准号:60776817
-
项目类别:联合基金项目
-
资助金额:20.0万元
-
批准年份:2007
-
负责人:李金林
-
依托单位:
供应链管理中的稳健型(Robust)策略分析和稳健型优化(Robust Optimization )方法研究
-
批准号:70601028
-
项目类别:青年科学基金项目
-
资助金额:7.0万元
-
批准年份:2006
-
负责人:王明征
-
依托单位:
心理紧张和应力影响下Robust语音识别方法研究
-
批准号:60085001
-
项目类别:专项基金项目
-
资助金额:14.0万元
-
批准年份:2000
-
负责人:韩纪庆
-
依托单位:
ROBUST语音识别方法的研究
-
批准号:69075008
-
项目类别:面上项目
-
资助金额:3.5万元
-
批准年份:1990
-
负责人:高雨青
-
依托单位:
改进型ROBUST序贯检测技术
-
批准号:68671030
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1986
-
负责人:刘有恒
-
依托单位: