课题基金 / 基金详情

Theory and Algorithms for Relation between Stochastic Control and Reinforcement Learning

Theory and Algorithms for Relation between Stochastic Control and Reinforcement Learning
随机控制与强化学习关系的理论和算法
批准号:
2741077
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
已结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
强化学习(RL)是机器学习中一个非常活跃的子领域。近年来,它在复杂决策问题领域的著名现实世界胜利包括玩完全信息棋盘游戏,如围棋和驾驶自动驾驶汽车。其基本思想是让算法逐渐学习接近最佳的策略,以在真实的时间内解决问题。学习是基于尝试和错误,以改善控制策略。它需要指定一个明确的目标得分函数的黑盒环境,其中包括环境响应的控制行动。然而,在许多情况下,这种反复试验的探索成本过高。RL中的探索-利用二分法通常通过添加需要作为目标函数的一部分被最小化的熵正则化项来减轻。最近,熵正则化RL问题和松弛连续时间随机控制之间的联系已经建立。特别是,线性二次(LQ)控制问题有优雅的解决方案,并能够近似更一般的问题。目标之一是设计和实现算法来解决这个问题,无论是LQ问题或一般控制问题,并提供相应的收敛定理。为此,Szpruch等人和Reisinger等人提出的算法是很好的例子,可以作为起点。在实际问题中,带跳的随机过程有着广泛的应用。因此,该项目的另一个目标是考虑更复杂的受控动力学,即跳跃过程,看看我们的算法是否可以解决这个更复杂的问题。Hernandel-Hernandez等人提出的鞅方法可以应用于该项目。在本研究中,我们将主要研究该问题的理论。当理论比较完善时,我们将尝试发展相应的算法,并为它们提供收敛性定理。有了这些理论和算法,我们将继续这个项目的最后一个目标:将算法应用于不同的现实世界问题,如金融交易,自动驾驶和棋盘游戏。RL在金融领域的许多应用尚未纳入连续时间松弛随机控制框架。据设想,在这个博士学位期间,理论进展将应用于定量金融问题(例如,投资组合交易的最优执行和衍生证券的风险管理),重点是算法的现实世界适用性。
英文摘要
Reinforcement learning (RL) is an extremely active subfield in machine learning. Its famous real-world triumphs in the area of complex decision making problems in recent years include playing perfect-information board games such as Go and driving autonomous vehicles. The basic idea is for an algorithm to learn gradually near-best strategies to solve a problem in real time. The learning is based on trial-and-error for improving control policies. It requires specifying an explicit objective score function on a black-box environment, which incorporates the environmental responses to the control actions. However, such trial-and-error exploration is in many situations prohibitively costly. The exploration-exploitation dichotomy in RL is typically mitigated by adding an entropy-regularisation term that needs to be minimised as part of the objective function. The connection between entropy-regularized RL problems and the relaxed continuous-time stochastic control has been established recently. In particular, Linear-Quadratic (LQ) control problem has elegant solutions and is able to approximate more general problems. One of the objective is to design and implement algorithms for solving this problem, either LQ problem or general control problem, and provide corresponding convergence theorems. To this end, algorithms proposed by Szpruch et al. and Reisinger et al. are good examples and could be a starting point. In real-world problems, stochastic process with jumps is widely applied. Thus, another objective of this project is to consider more complex controlled dynamics, i.e. processes with jump, and see whether our algorithms could tackle this more complex problem. The martingale approach proposed by Hernandez-Hernandez et al. could be applied to this project. In this research, we will mainly study the theory of the problem. When the theory is relatively complete, we will try to develop corresponding algorithms and provide convergence theorems for them. With these theory and algorithms, we will proceed to the last objective of this project: applying the algorithms to different real-world problem, such as financial trading, autonomous driving and board games. Many applications of RL in finance are yet to be incorporated in the continuous-time relaxed stochastic control framework. It is envisaged that, during this PhD, theoretical advances will be applied to problems in quantitative finance (e.g. optimalexecution of portfolio transactions and the risk-management of derivative securities) with emphasis on real-worlds applicability of the algorithms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金