Optimal Stopping with Unknown Gain Function
Optimal Stopping with Unknown Gain Function
批准号:
2585636
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
在最优停止问题中,增益函数的形式往往是未知的。一种可能的解决方案是采用模仿学习的方法,从专家的演示中推断增益函数。模仿学习被认为是强化学习的一个分支,最近被证明是解决最优停止和/或最优控制问题的一个有用的方法。该项目的目的是开发一个数学支持的框架来解决具有未知增益函数的最优停止问题(逆最优停止)。目标包括:-建立一般以及逆最优停止和最优控制问题的强化学习公式;-识别现有最优停止/最优控制问题方法的主要缺陷;-开发强化学习算法以高效和有效地解决逆最优停止问题并为其建立数学保证;- 将所开发的算法应用于自动驾驶汽车控制中。据我们所知,关于具有未知增益函数的最优停止问题的文献有限,并且该领域的现有研究主要涵盖了现有算法的应用,而没有深入研究。深入的数学证明和算法的收敛性和稳定性的保证。该项目符合EPSRC涵盖数学科学领域的职权范围,并与研究理事会的人工智能和机器人团队进行的活动密切相关。
英文摘要
In the optimal stopping problems the form of the gain function is often unknown. One of the possible solutions to this is to employ the approach of imitation learning to infer the gain function from the expert's demonstrations. Imitation learning is considered as a branch of Reinforcement Learning which recently proved to be a useful in solving the optimal stopping and/or optimal control problems. The aim of the Project is to develop a mathematically backed framework to solving the optimal stopping problems with an unknown gain function (inverse optimal stopping). The objectives include- Establishing a Reinforcement Learning formulation of the general as well as the inverse optimal stopping and optimal control problems;- Identifying the main pitfalls of the existing approaches to the optimal stopping/optimal control problems;- Developing a Reinforcement Learning algorithm to efficiently and effectively solve the inverse optimal stopping problems and establishing mathematical guarantees for it;- Creating an application of the developed algorithms with a potential to be used in the autonomous vehicles control.To the best of our knowledge there is a limited literature available on the topic of optimal stopping problems with an unknown gain function and the existing research in the area mainly covers the applications of the existing algorithms without in-depth mathematical proofs and guarantees of the algorithm's convergence and stability.The project aligns with the EPSRC remit covering the area of Mathematical Sciences and is closely related to the activities conducted by the AI and Robotics team of the Research Council.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金