课题基金 / 基金详情

Optimal Stopping with Unknown Gain Function

Optimal Stopping with Unknown Gain Function
未知增益函数的最佳停止
批准号:
2585636
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在最优停止问题中,增益函数的形式通常是未知的。一种可能的解决方案是采用模仿学习的方法,从专家的演示中推断增益函数。模仿学习被认为是强化学习的一个分支,最近被证明在解决最优停止和/或最优控制问题方面很有用。该项目的目的是开发一个数学支持的框架来解决未知增益函数(逆最优停止)的最优停止问题。目标包括-建立一般以及逆最优停止和最优控制问题的强化学习公式;-找出解决最佳停止/最佳控制问题的现有方法的主要缺陷;-开发强化学习算法,高效有效地解决逆最优停车问题,并建立数学保证;-将开发的算法应用于自动驾驶汽车控制。据我们所知,关于增益函数未知的最优停止问题的文献有限,现有的研究主要涵盖了现有算法的应用,没有深入的数学证明,也没有保证算法的收敛性和稳定性。该项目符合EPSRC在数学科学领域的职权范围,并与研究委员会人工智能和机器人团队的活动密切相关。
英文摘要
In the optimal stopping problems the form of the gain function is often unknown. One of the possible solutions to this is to employ the approach of imitation learning to infer the gain function from the expert's demonstrations. Imitation learning is considered as a branch of Reinforcement Learning which recently proved to be a useful in solving the optimal stopping and/or optimal control problems. The aim of the Project is to develop a mathematically backed framework to solving the optimal stopping problems with an unknown gain function (inverse optimal stopping). The objectives include- Establishing a Reinforcement Learning formulation of the general as well as the inverse optimal stopping and optimal control problems;- Identifying the main pitfalls of the existing approaches to the optimal stopping/optimal control problems;- Developing a Reinforcement Learning algorithm to efficiently and effectively solve the inverse optimal stopping problems and establishing mathematical guarantees for it;- Creating an application of the developed algorithms with a potential to be used in the autonomous vehicles control.To the best of our knowledge there is a limited literature available on the topic of optimal stopping problems with an unknown gain function and the existing research in the area mainly covers the applications of the existing algorithms without in-depth mathematical proofs and guarantees of the algorithm's convergence and stability.The project aligns with the EPSRC remit covering the area of Mathematical Sciences and is closely related to the activities conducted by the AI and Robotics team of the Research Council.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金