Evaluating the Stability of Non-Adaptive Trading in Continuous Double Auctions: A Reinforcement Learning Approach

Evaluating the Stability of Non-Adaptive Trading in Continuous Double Auctions: A Reinforcement Learning Approach
复制标题

评估连续双重拍卖中非自适应交易的稳定性:一种强化学习方法

DOI:
--
复制
发表时间:
2018
期刊:
AAAI Workshops
影响因子:
--
通讯作者:
Michael P. Wellman
Michael P. Wellman
中科院分区:
--
文献类型:
--
作者:
Mason Wright;Michael P. Wellman

文献摘要

被引文献

相似文献

连续双向拍卖(CDA)是现代证券市场的主导机制。许多基于代理的CDA环境分析依赖于简单的非自适应交易策略,如零智能(ZI),这(正如他们的名字所暗示的)是非常有限的。我们研究的可行性,这种依赖,通过实证博弈论分析在一个合理的市场环境。具体来说,我们评估了一小部分ZI交易者在更大的政策空间上应用强化学习(RL)发现的策略时,平衡的策略稳定性。RL确实可以通过对交易执行的可能性或当前出价和要价的可接受性的信号进行调节,找到与ZI交易者均衡的有益偏离。然而,经验观察到,校准良好的ZI政策所获得的盈余几乎与适应性策略所能获得的盈余一样大,尽管它们具有更大的政策空间。我们的研究结果普遍支持使用平衡ZI交易者在CDA研究。
The continuous double auction (CDA) is the predominant mechanism in modern securities markets. Many agent-based analyses of CDA environments rely on simple non-adaptive trading strategies like Zero Intelligence (ZI), which (as their name suggests) are quite limited. We examine the viability of this reliance through empirical game-theoretic analysis in a plausible market environment. Specifically, we evaluate the strategic stability of equilibria defined over a small set of ZI traders with respect to strategies found by reinforcement learning (RL) applied over a much larger policy space. RL can indeed find beneficial deviations from equilibria of ZI traders, by conditioning on signals of the likelihood a trade will execute or the favorability of the current bid and ask. Nevertheless, the surplus earned by well-calibrated ZI policies is empirically observed to be nearly as great as what the adaptive strategies can earn, despite their much more expressive policy space. Our findings generally support the use of equilibrated ZI traders in CDA studies.