RankGen: Improving Text Generation with Large Ranking Models

RankGen: Improving Text Generation with Large Ranking Models
复制标题

DOI:
10.48550/arxiv.2205.09726
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Kalpesh Krishna;Yapei Chang;J. Wieting;Mohit Iyyer
Kalpesh Krishna;Yapei Chang;J. Wieting;Mohit Iyyer
中科院分区:
其他
文献类型:
--
作者:
Kalpesh Krishna;Yapei Chang;J. Wieting;Mohit Iyyer

文献摘要

被引文献

相似文献

给定输入序列(或前缀),现代语言模型通常将高概率分配给与前缀重复、不连贯或无关的输出序列;因此,模型生成的文本也包含这样的人工产物。为了解决这些问题,我们提出了RankGen,这是一个用于英语的1.2B参数编码器模型,它给给定前缀的模型代打分。RankGen可以灵活地合并为波束搜索中的评分函数,并用于从任何预先训练的语言模型进行解码。我们使用大规模的对比学习来训练RankGen将一个前缀映射到紧随其后的基本事实序列,而远离两种类型的否定:(1)来自与前缀相同的文档的随机序列,以及(2)由大型语言模型生成的以前缀为条件的序列。在四个不同的语言模型(345M-11B参数)和两个域上的实验表明,RankGen在自动指标(85.0vs77.3 Muve)和人类对英语作家的评估(74.5%的人类偏好)上都显著优于NITUS、TOP-K和Typical Samples等解码算法。分析表明,与基线相比,RankGen产出与前缀更相关,并提高了连续性和一致性。我们发布了我们的模型检查点、代码和人类偏好数据,并提供了解释,以便于未来的研究。
Given an input sequence (or prefix), modern language models often assign high probabilities to output sequences that are repetitive, incoherent, or irrelevant to the prefix; as such, model-generated text also contains such artifacts. To address these issues we present RankGen, a 1.2B parameter encoder model for English that scores model generations given a prefix. RankGen can be flexibly incorporated as a scoring function in beam search and used to decode from any pretrained language model. We train RankGen using large-scale contrastive learning to map a prefix close to the ground-truth sequence that follows it and far away from two types of negatives: (1) random sequences from the same document as the prefix, and (2) sequences generated from a large language model conditioned on the prefix. Experiments across four different language models (345M-11B parameters) and two domains show that RankGen significantly outperforms decoding algorithms like nucleus, top-k, and typical sampling on both automatic metrics (85.0 vs 77.3 MAUVE) as well as human evaluations with English writers (74.5% human preference over nucleus sampling). Analysis reveals that RankGen outputs are more relevant to the prefix and improve continuity and coherence compared to baselines. We release our model checkpoints, code, and human preference data with explanations to facilitate future research.