课题基金 / 基金详情

Theory and Algorithms for Relation between Stochastic Control and Reinforcement Learning

Theory and Algorithms for Relation between Stochastic Control and Reinforcement Learning
随机控制与强化学习关系的理论和算法
批准号:
2741077
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
已结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Reinforcement learning (RL) is an extremely active subfield in machine learning. Its famous real-world triumphs in the area of complex decision making problems in recent years include playing perfect-information board games such as Go and driving autonomous vehicles. The basic idea is for an algorithm to learn gradually near-best strategies to solve a problem in real time. The learning is based on trial-and-error for improving control policies. It requires specifying an explicit objective score function on a black-box environment, which incorporates the environmental responses to the control actions. However, such trial-and-error exploration is in many situations prohibitively costly. The exploration-exploitation dichotomy in RL is typically mitigated by adding an entropy-regularisation term that needs to be minimised as part of the objective function. The connection between entropy-regularized RL problems and the relaxed continuous-time stochastic control has been established recently. In particular, Linear-Quadratic (LQ) control problem has elegant solutions and is able to approximate more general problems. One of the objective is to design and implement algorithms for solving this problem, either LQ problem or general control problem, and provide corresponding convergence theorems. To this end, algorithms proposed by Szpruch et al. and Reisinger et al. are good examples and could be a starting point. In real-world problems, stochastic process with jumps is widely applied. Thus, another objective of this project is to consider more complex controlled dynamics, i.e. processes with jump, and see whether our algorithms could tackle this more complex problem. The martingale approach proposed by Hernandez-Hernandez et al. could be applied to this project. In this research, we will mainly study the theory of the problem. When the theory is relatively complete, we will try to develop corresponding algorithms and provide convergence theorems for them. With these theory and algorithms, we will proceed to the last objective of this project: applying the algorithms to different real-world problem, such as financial trading, autonomous driving and board games. Many applications of RL in finance are yet to be incorporated in the continuous-time relaxed stochastic control framework. It is envisaged that, during this PhD, theoretical advances will be applied to problems in quantitative finance (e.g. optimalexecution of portfolio transactions and the risk-management of derivative securities) with emphasis on real-worlds applicability of the algorithms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金