Trading Off Diversity and Quality in Natural Language Generation

Trading Off Diversity and Quality in Natural Language Generation
复制标题

DOI:
--
复制
发表时间:
2020-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Hugh Zhang;Daniel Duckworth;Daphne Ippolito;Arvind Neelakantan
Hugh Zhang;Daniel Duckworth;Daphne Ippolito;Arvind Neelakantan
中科院分区:
其他
文献类型:
--
作者:
Hugh Zhang;Daniel Duckworth;Daphne Ippolito;Arvind Neelakantan

文献摘要

被引文献

相似文献

对于开放式语言生成任务,如讲故事或对话,选择正确的解码算法对于控制生成质量和多样性之间的权衡至关重要。然而,目前对于哪种解码程序是最好的,甚至对于比较它们的标准,还没有达成共识。在本文中,我们投解码作为响应质量和多样性之间的权衡,我们执行的第一个大规模的评估解码方法沿着整个质量多样性频谱。我们的实验证实了似然陷阱的存在:高似然序列的质量往往出奇地低,这是一种违反直觉的观察。我们还发现,当多样性是一个优先事项时,所有方法的表现相似,但当质量被视为更重要时,细胞核采样(Holtzman等人,2019)优于所有其他评估的解码算法。
For open-ended language generation tasks such as storytelling or dialogue, choosing the right decoding algorithm is vital for controlling the tradeoff between generation quality and diversity. However, there presently exists no consensus on which decoding procedure is best or even the criteria by which to compare them. In this paper, we cast decoding as a tradeoff between response quality and diversity, and we perform the first large-scale evaluation of decoding methods along the entire quality-diversity spectrum. Our experiments confirm the existence of the likelihood trap: the counter-intuitive observation that high likelihood sequences are often surprisingly low quality. We also find that when diversity is a priority, all methods perform similarly, but when quality is viewed as more important, nucleus sampling (Holtzman et al., 2019) outperforms all other evaluated decoding algorithms.