Data-Driven Sentence Simplification: Survey and Benchmark

Data-Driven Sentence Simplification: Survey and Benchmark
复制标题

DOI:
10.1162/coli_a_00370
复制
发表时间:
2020-03-01
影响因子:
9.3
通讯作者:
Specia, Lucia
Specia, Lucia
中科院分区:
计算机科学3区
文献类型:
--
作者:
Alva-Manchego, Fernando;Scarton, Carolina;Specia, Lucia

文献摘要

被引文献

相似文献

句子简化的目的是对句子进行修饰,使其更易于阅读和理解。为此,可以执行若干重写转换,例如替换、重新排序和拆分。在执行这些转换的同时保持句子的语法,保留其主要思想,并生成更简单的输出,这是一个具有挑战性的问题,而且远未得到解决。在本文中,我们综述了对源句的研究,重点是学习如何简化使用英语中对齐的原始-简句对的语料库的方法,这是当今占主导地位的范例。我们还包括关于共同数据集的不同方法的基准,以便对它们进行比较,并突出它们的优点和局限性。我们预计,这项调查将作为对这项任务感兴趣的研究人员的一个起点,并有助于为未来的发展激发新的想法。
Sentence Simplification (SS) aims to modify a sentence in order to make it easier to read and understand. In order to do so, several rewriting transformations can be performed such as replacement, reordering, and splitting. Executing these transformations while keeping sentences grammatical, preserving their main idea, and generating simpler output, is a challenging and still far from solved problem. In this article, we survey research on SS, focusing on approaches that attempt to learn how to simplify using corpora of aligned original-simplified sentence pairs in English, which is the dominant paradigm nowadays. We also include a benchmark of different approaches on common data sets so as to compare them and highlight their strengths and limitations. We expect that this survey will serve as a starting point for researchers interested in the task and help spark new ideas for future developments.