Double Deep Q-Learning for Optimal Execution

Double Deep Q-Learning for Optimal Execution
复制标题

双深度 Q-Learning 实现最佳执行

DOI:
10.1080/1350486x.2022.2077783
复制
发表时间:
2018
影响因子:
--
通讯作者:
S. Jaimungal
S. Jaimungal
中科院分区:
--
文献类型:
--
作者:
Brian Ning;Franco Ho Ting Ling;S. Jaimungal

文献摘要

被引文献

相似文献

摘要最优交易执行是所有交易者都面临的重要问题。许多研究最优执行使用严格的模型假设,并应用连续时间随机控制来解决这些问题。在这里,我们采用了一种无模型的方法,并开发了一种深度Q学习的变体来估计交易者的最佳行为。该模型是一个完全连接的神经网络,使用经验重放和双DQN进行训练,输入特征由限价订单簿的当前状态,其他交易信号和可用的执行操作给出,而输出是Q值函数,估计任意操作下的未来回报。我们将我们的模型应用于9只不同的股票,发现它在大多数股票上的表现优于标准的基准方法,这些方法使用的指标包括:(i)平均和中位数的表现,(ii)表现出色的概率,(iii)损益比。
ABSTRACT Optimal trade execution is an important problem faced by essentially all traders. Much research into optimal execution uses stringent model assumptions and applies continuous time stochastic control to solve them. Here, we instead take a model free approach and develop a variation of Deep Q-Learning to estimate the optimal actions of a trader. The model is a fully connected Neural Network trained using Experience Replay and Double DQN with input features given by the current state of the limit order book, other trading signals, and available execution actions, while the output is the Q-value function estimating the future rewards under an arbitrary action. We apply our model to nine different stocks and find that it outperforms the standard benchmark approach on most stocks using the measures of (i) mean and median out-performance, (ii) probability of out-performance, and (iii) gain-loss ratios.