Reinforcement Learning in Multi-Party Trading Dialog
Reinforcement Learning in Multi-Party Trading Dialog
复制标题
DOI:
10.18653/v1/w15-4605
复制
发表时间:
2015-09
期刊:
影响因子:
--
通讯作者:
Takuya Hiraoka;Kallirroi Georgila;E. Nouri;D. Traum;Satoshi Nakamura
中科院分区:
文献类型:
--
作者:
Takuya Hiraoka;Kallirroi Georgila;E. Nouri;D. Traum;Satoshi Nakamura
In this paper, we apply reinforcement learning (RL) to a multi-party trading scenario where the dialog system (learner) trades with one, two, or three other agents. We experiment with different RL algorithms and reward functions. The negotiation strategy of the learner is learned through simulated dialog with trader simulators. In our experiments, we evaluate how the performance of the learner varies depending on the RL algorithm used and the number of traders. Our results show that (1) even in simple multi-party trading dialog tasks, learning an effective negotiation policy is a very hard problem; and (2) the use of neural fitted Q iteration combined with an incremental reward function produces negotiation policies as effective or even better than the policies of two strong hand-crafted baselines.