Reliability and Learnability of Human Bandit Feedback for Sequence-to-Sequence Reinforcement Learning

Reliability and Learnability of Human Bandit Feedback for Sequence-to-Sequence Reinforcement Learning
复制标题

DOI:
10.18653/v1/p18-1165
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Julia Kreutzer;Joshua Uyheng;S. Riezler
Julia Kreutzer;Joshua Uyheng;S. Riezler
中科院分区:
其他
文献类型:
--
作者:
Julia Kreutzer;Joshua Uyheng;S. Riezler

文献摘要

相似文献

我们提出了一个研究强化学习(RL)从人类强盗反馈序列到序列的学习,以强盗神经机器翻译(NMT)的任务为例。我们研究了人类强盗反馈的可靠性,并分析了可靠性对奖励估计的可学习性的影响,以及奖励估计的质量对整个RL任务的影响。我们对基数(5点评分)和序数(成对偏好)反馈的分析表明,它们的注释者内和注释者间α一致性是相当的。标准化基数反馈的可靠性最好,基数反馈也最容易学习和推广。最后,通过将基于回归的奖励估计器(在800个翻译的基数反馈上训练)集成到NMT的RL中,可以获得超过1个BLEU的改进。这表明,即使是从少量相当可靠的人类反馈中,RL也是可能的,这表明了大规模应用的巨大潜力。
We present a study on reinforcement learning (RL) from human bandit feedback for sequence-to-sequence learning, exemplified by the task of bandit neural machine translation (NMT). We investigate the reliability of human bandit feedback, and analyze the influence of reliability on the learnability of a reward estimator, and the effect of the quality of reward estimates on the overall RL task. Our analysis of cardinal (5-point ratings) and ordinal (pairwise preferences) feedback shows that their intra- and inter-annotator α-agreement is comparable. Best reliability is obtained for standardized cardinal feedback, and cardinal feedback is also easiest to learn and generalize from. Finally, improvements of over 1 BLEU can be obtained by integrating a regression-based reward estimator trained on cardinal feedback for 800 translations into RL for NMT. This shows that RL is possible even from small amounts of fairly reliable human feedback, pointing to a great potential for applications at larger scale.