Reinforcement Learning in Multi-Party Trading Dialog

Reinforcement Learning in Multi-Party Trading Dialog
复制标题

DOI:
10.18653/v1/w15-4605
复制
发表时间:
2015-09
期刊:
--
影响因子:
--
通讯作者:
Takuya Hiraoka;Kallirroi Georgila;E. Nouri;D. Traum;Satoshi Nakamura
Takuya Hiraoka;Kallirroi Georgila;E. Nouri;D. Traum;Satoshi Nakamura
中科院分区:
其他
文献类型:
--
作者:
Takuya Hiraoka;Kallirroi Georgila;E. Nouri;D. Traum;Satoshi Nakamura

文献摘要

被引文献

相似文献

在本文中,我们将强化学习(RL)应用于多方交易场景,其中对话系统(学习者)与一个,两个或三个其他代理进行交易。我们使用不同的RL算法和奖励函数进行实验。学习者的谈判策略是通过与交易模拟器的模拟对话来学习的。在我们的实验中,我们评估了学习者的表现如何根据所使用的RL算法和交易者的数量而变化。我们的研究结果表明:(1)即使在简单的多方交易对话任务中,学习有效的谈判策略也是一个非常困难的问题;(2)使用神经拟合Q迭代结合增量奖励函数产生的谈判策略与两个强大的手工基线的策略一样有效,甚至更好。
In this paper, we apply reinforcement learning (RL) to a multi-party trading scenario where the dialog system (learner) trades with one, two, or three other agents. We experiment with different RL algorithms and reward functions. The negotiation strategy of the learner is learned through simulated dialog with trader simulators. In our experiments, we evaluate how the performance of the learner varies depending on the RL algorithm used and the number of traders. Our results show that (1) even in simple multi-party trading dialog tasks, learning an effective negotiation policy is a very hard problem; and (2) the use of neural fitted Q iteration combined with an incremental reward function produces negotiation policies as effective or even better than the policies of two strong hand-crafted baselines.