BiSECT: Learning to Split and Rephrase Sentences with Bitexts

BiSECT: Learning to Split and Rephrase Sentences with Bitexts
复制标题

DOI:
10.18653/v1/2021.emnlp-main.500
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Joongwon Kim;Mounica Maddela;Reno Kriz;Wei Xu;Chris Callison-Burch
Joongwon Kim;Mounica Maddela;Reno Kriz;Wei Xu;Chris Callison-Burch
中科院分区:
其他
文献类型:
--
作者:
Joongwon Kim;Mounica Maddela;Reno Kriz;Wei Xu;Chris Callison-Burch

文献摘要

相似文献

NLP 应用(例如句子简化)中的一项重要任务是能够将长而复杂的句子拆分为较短的句子,并根据需要重新措辞。我们为这个“分割和改写”任务引入了一个新颖的数据集和一个新模型。我们的 BiSECT 训练数据由 100 万个长英语句子和较短的、含义相当的英语句子组成。我们通过在双语平行语料库中提取 1-2 个句子对齐,然后使用机器翻译将语料库双方转换为相同语言来获得这些信息。 BiSECT 包含比之前的 Split 和 Rephrase 语料库更高质量的训练示例,句子分割需要更重大的修改。我们对语料库中的示例进行分类,并在一个新颖的模型中使用这些类别,该模型允许我们针对要分割和编辑的输入句子的特定区域。此外,我们还表明,在 BiSECT 上训练的模型可以执行更广泛的分割操作,并改进了之前自动和人工评估中最先进的方法。
An important task in NLP applications such as sentence simplification is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary. We introduce a novel dataset and a new model for this ‘split and rephrase’ task. Our BiSECT training data consists of 1 million long English sentences paired with shorter, meaning-equivalent English sentences. We obtain these by extracting 1-2 sentence alignments in bilingual parallel corpora and then using machine translation to convert both sides of the corpus into the same language. BiSECT contains higher quality training examples than the previous Split and Rephrase corpora, with sentence splits that require more significant modifications. We categorize examples in our corpus and use these categories in a novel model that allows us to target specific regions of the input sentence to be split and edited. Moreover, we show that models trained on BiSECT can perform a wider variety of split operations and improve upon previous state-of-the-art approaches in automatic and human evaluations.