Developing robust and scalable reinforcement learning algorithms
Developing robust and scalable reinforcement learning algorithms
批准号:
2740739
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Reinforcement learning (RL) represents a powerful paradigm for applying machine learning to complex decision making tasks. However, current RL algorithms are brittle and have numerous failure mechanisms, making them difficult to apply to beyond the synthetic simple environments used in research. Over my PhD I aim to help in developing robust and scalable RL algorithms, such that RL can be more easily applied to large, complex, real-world problems.One promising approach to this is offline RL; a broad set of methods which utilize pre-existing datasets to remove the need for the online data collection required by standard RL methods. Use of large-scale datasets has been the driver of the transformative recent advances in supervised learning, and offline RL presents a way of enabling similar scaling progress in RL. In addition to this, offline RL is a practical choice in many real-world domains where online data collection is costly or poses safety risks, such as in robotics or healthcare. The offline approach, however, introduces the extra problem of needing to evaluate effectiveness of actions or behaviour not covered in the dataset. This additional challenge has been shown to cause existing online RL algorithms to fail, and hence represents a significant barrier to robust, scalable offline RL algorithms.In my recent project we investigated this issue and identified the "edge-of-reach" failure mechanism in offline model-based reinforcement learning. Based on these insights we were able propose a "value uncertainty-aware" approach which resolves this issue. As an extension to this, I plan to investigate whether meta-learning can be used to discover an offline RL algorithm that deals with value uncertainty even more effectively.Another barrier to large-scale RL is the phenomenon of "plasticity loss," whereby the inherent non-stationarity in RL training has been shown to cause deep RL agents to gradually lose the ability to learn from new data. I am currently investigating the underlying mechanisms behind this phenomenon, with the hope that this will inform development of more effective mitigation techniques.Finally, the recent advances in large language models present huge opportunity for using language as a tool for instilling real-world knowledge and inductive biases into RL agents. The most effective approach for this integration is an open question, and I plan to work on exploring this. In summary, RL has huge potential which is not yet realized by current approaches, and during my PhD I aim to contribute towards developing RL algorithms which are robust and scalable and are effective when applied to complex real-world problems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
半定松弛与非凸二次约束二次规划研究
-
批准号:11271243
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2012
-
负责人:王燕军
-
依托单位:
基于复合编码脉冲串的水下主动隐蔽性探测新方法研究
-
批准号:61271414
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2012
-
负责人:冯西安
-
依托单位:
民航客运网络收益管理若干问题的研究
-
批准号:60776817
-
项目类别:联合基金项目
-
资助金额:20.0万元
-
批准年份:2007
-
负责人:李金林
-
依托单位:
供应链管理中的稳健型(Robust)策略分析和稳健型优化(Robust Optimization )方法研究
-
批准号:70601028
-
项目类别:青年科学基金项目
-
资助金额:7.0万元
-
批准年份:2006
-
负责人:王明征
-
依托单位:
心理紧张和应力影响下Robust语音识别方法研究
-
批准号:60085001
-
项目类别:专项基金项目
-
资助金额:14.0万元
-
批准年份:2000
-
负责人:韩纪庆
-
依托单位:
ROBUST语音识别方法的研究
-
批准号:69075008
-
项目类别:面上项目
-
资助金额:3.5万元
-
批准年份:1990
-
负责人:高雨青
-
依托单位:
改进型ROBUST序贯检测技术
-
批准号:68671030
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1986
-
负责人:刘有恒
-
依托单位: