Random Policy Search and Stochastic control
Random Policy Search and Stochastic control
批准号:
2435718
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The paradigm of model free reinforcement learning has proved successful in a multitude of tasks in automated decision making. This paradigm assumes limited information about the environment it operates in and tries to learn good policies by interacting with its environment over time. The quality of a policy is judged by its performance on an appropriately chosen objective function. Usually, the learning algorithms rely on estimating either a value function or directly the policy or a mixture of both. A specific challenge is the learning of policies in environments with continuous action spaces for which computer scientists have introduced several simulation-based benchmark problems (e.g. MuJoCo). A simple algorithm that produced remarkably good results in these benchmarks is the random policy search algorithm with linear parametrization. For this algorithm, our goal is to find a good policy directly by learning parametrized policies linear in the state instead of searching the entire function space of admissible policies. The fitting is done by approximating the gradient of the value function in the directions of the parameter with a Monte-Carlo scheme. Given the gradient, a gradient ascent step is performed. This procedure is repeated until a suitable policy is found and only requires a sampling oracle.In spite of the simplicity of the algorithm, little is known about its theoretical properties. A good starting point for an analysis is stochastic control theory. Already, authors used these tools to establish convergence to the optimal policy in cases where the algorithm is applied to the linear quadratic regulator (LQR). We would like to build on the same approach to establish the behaviour of the algorithm in other control frameworks and deepen the understanding of the interplay between representation and optimization. This research is of interest since it potentially improves algorithm selection for practical problems and gives new insights in deeper aspects of learning.For the first part, we apply the random policy search algorithm to a time-continuous optimal queueing problem that has a direct application to optimal execution in limit order book markets. For this problem and a specified linear policy on discretized time intervals, we can give the closed form expression for the value function and its gradient. We hope to find a parameter region for which the algorithm is guaranteed to converge to an optimal time-discrete policy, if initialised within the region. For that purpose, one could establish a gradient dominance condition for a gradient-descent algorithm with the known gradient and then generalize to a situation where the gradient is approximated. Furthermore, since the value function for the optimal time-continuous policy can be calculated as well we hope to make a statement about discretization errors. Finally, we want to apply the algorithm to real world limit order book data and extend it to a two sided queue corresponding to the market maker problem. I work with Roel Oomen from Deutsche Bank on this project.Further, we would like to better understand the capability of the algorithm to learn an appropriate state representation while simultaneously optimizing. As a first instructive problem, we could consider the LQR case where the appropriate state can be almost written as a combination of bases functions where the coefficients are assumed to have a low dimensional structure. The optimal result in this framework would establish a model selection property for a proximal gradient descent algorithm. Finally, we hope to investigate the role of state representation in an application of this algorithm in the context of backward stochastic differential equation solvers. This project falls within the EPSRC Artificial intelligence technologies research area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
The Heterogenous Impact of Monetary Policy on Firms' Risk and Fundamentals
-
批准号:--
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:潘军
-
依托单位:
Financial Constraints in China
and Their Policy Implications
-
批准号:--
-
项目类别:外国优秀青年学 者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:Jake Zhao
-
依托单位: