Sentence-Level Content Planning and Style Specification for Neural Text Generation

Sentence-Level Content Planning and Style Specification for Neural Text Generation
复制标题

DOI:
10.18653/v1/d19-1055
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Xinyu Hua;Lu Wang
Xinyu Hua;Lu Wang
中科院分区:
其他
文献类型:
--
作者:
Xinyu Hua;Lu Wang

文献摘要

相似文献

构建有效的文本生成系统需要三个关键组成部分:内容选择,文本规划和表面实现,传统上它们是作为单独的问题来处理的。最近的一体式神经生成模型取得了令人印象深刻的进展,但它们经常产生不连贯且不忠实于输入的输出。为了解决这些问题,我们提出了一个端到端的训练两步生成模型,其中一个文档级内容规划器首先决定要覆盖的关键短语以及所需的语言风格,然后是一个生成相关和连贯文本的表面实现解码器。对于实验,我们考虑了三个任务,从不同的主题和不同的语言风格的域:有说服力的论点建设从Reddit,段落生成的正常和简单版本的维基百科,和摘要生成的科学文章。自动评估表明,我们的系统可以显着优于竞争对手的比较。与不考虑语言风格的变体相比,人类法官进一步将我们的系统生成的文本评为更流畅和正确。
Building effective text generation systems requires three critical components: content selection, text planning, and surface realization, and traditionally they are tackled as separate problems. Recent all-in-one style neural generation models have made impressive progress, yet they often produce outputs that are incoherent and unfaithful to the input. To address these issues, we present an end-to-end trained two-step generation model, where a sentence-level content planner first decides on the keyphrases to cover as well as a desired language style, followed by a surface realization decoder that generates relevant and coherent text. For experiments, we consider three tasks from domains with diverse topics and varying language styles: persuasive argument construction from Reddit, paragraph generation for normal and simple versions of Wikipedia, and abstract generation for scientific articles. Automatic evaluation shows that our system can significantly outperform competitive comparisons. Human judges further rate our system generated text as more fluent and correct, compared to the generations by its variants that do not consider language style.