Expert-based reward function training: the novel method to train sequence generators

Expert-based reward function training: the novel method to train sequence generators
复制标题

DOI:
--
复制
发表时间:
2018-04
期刊:
--
影响因子:
--
通讯作者:
Joji Toyama;Yusuke Iwasawa;Kotaro Nakayama;Y. Matsuo
Joji Toyama;Yusuke Iwasawa;Kotaro Nakayama;Y. Matsuo
中科院分区:
其他
文献类型:
--
作者:
Joji Toyama;Yusuke Iwasawa;Kotaro Nakayama;Y. Matsuo

文献摘要

相似文献

序列发生器的训练方法与策略梯度的组合在本文中表现出良好的性能,我们提出了基于专家的奖励功能训练:训练序列生成器的新方法。基于奖励功能的培训不利用GAN的框架,我们的模型表现优于Seqgan和强大的基线Rankgan。
The training methods of sequence generator with a combination of GAN and policy gradient has shown good performance. In this paper, we propose expert-based reward function training: the novel method to train sequence generator. Different from previous studies of sequence generation, expert-based reward function training does not utilize GAN’s framework. Still, our model outperforms SeqGAN and a strong baseline, RankGAN.