Generating Sequences by Learning to Self-Correct

Generating Sequences by Learning to Self-Correct
复制标题

DOI:
10.48550/arxiv.2211.00053
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Welleck;Ximing Lu;Peter West;Faeze Brahman;T. Shen;Daniel Khashabi;Yejin Choi
S. Welleck;Ximing Lu;Peter West;Faeze Brahman;T. Shen;Daniel Khashabi;Yejin Choi
中科院分区:
其他
文献类型:
--
作者:
S. Welleck;Ximing Lu;Peter West;Faeze Brahman;T. Shen;Daniel Khashabi;Yejin Choi

文献摘要

被引文献

相似文献

序列生成应用需要满足语义约束,例如确保程序正确、使用特定关键词或避免不良内容。语言模型,无论是经过微调的还是通过少量示例提示的,经常违反这些约束,并且缺乏一种迭代修改其输出的机制。此外,一些强大的语言模型规模极大或难以获取,这使得针对特定任务进行调整而更新其参数即使不是不可行,也是低效的。我们提出了自我修正(Self - Correction)方法,该方法将一个不完美的基础生成器(一个现成的语言模型或有监督的序列到序列模型)与一个单独的修正器分离,这个修正器学习迭代地修正不完美的生成内容。为了训练修正器,我们提出了一种在线训练过程,该过程可以使用对中间不完美生成内容的标量反馈或自然语言反馈。我们表明,自我修正在三个不同的生成任务——数学程序合成、词汇受限生成和毒性控制——中改进了基础生成器,即使修正器比基础生成器小得多。
Sequence generation applications require satisfying semantic constraints, such as ensuring that programs are correct, using certain keywords, or avoiding undesirable content. Language models, whether fine-tuned or prompted with few-shot demonstrations, frequently violate these constraints, and lack a mechanism to iteratively revise their outputs. Moreover, some powerful language models are of extreme scale or inaccessible, making it inefficient, if not infeasible, to update their parameters for task-specific adaptation. We present Self-Correction, an approach that decouples an imperfect base generator (an off-the-shelf language model or supervised sequence-to-sequence model) from a separate corrector that learns to iteratively correct imperfect generations. To train the corrector, we propose an online training procedure that can use either scalar or natural language feedback on intermediate imperfect generations. We show that Self-Correction improves upon the base generator in three diverse generation tasks - mathematical program synthesis, lexically-constrained generation, and toxicity control - even when the corrector is much smaller than the base generator.