Discourse Level Factors for Sentence Deletion in Text Simplification

Discourse Level Factors for Sentence Deletion in Text Simplification
复制标题

DOI:
10.1609/aaai.v34i05.6520
复制
发表时间:
2019-11
期刊:
--
影响因子:
--
通讯作者:
Yang Zhong;Chao Jiang;W. Xu;Junyi Jessy Li
Yang Zhong;Chao Jiang;W. Xu;Junyi Jessy Li
中科院分区:
其他
文献类型:
--
作者:
Yang Zhong;Chao Jiang;W. Xu;Junyi Jessy Li

文献摘要

相似文献

本文提出了一项数据驱动的研究,重点是在一个大型英语文本简化语料库上分析和预测句子删除--这是文档简化中一种普遍但研究较少的现象。我们使用一个新的人工标注的句子对齐语料库,考察了与句子删除相关的各种文档和语篇因素。我们发现,专业编辑使用不同的策略来满足中小学的可读性标准。为了预测句子在简化到一定程度时是否会被删除,我们利用自动对齐的数据来训练分类模型。根据我们人工标注的数据进行评估,我们最好的模型在小学和中学水平上的F1分数分别达到了65.2和59.7。我们发现,语篇层面的因素导致了为了简化而预测句子删除这一具有挑战性的任务。
This paper presents a data-driven study focusing on analyzing and predicting sentence deletion — a prevalent but understudied phenomenon in document simplification — on a large English text simplification corpus. We inspect various document and discourse factors associated with sentence deletion, using a new manually annotated sentence alignment corpus we collected. We reveal that professional editors utilize different strategies to meet readability standards of elementary and middle schools. To predict whether a sentence will be deleted during simplification to a certain level, we harness automatically aligned data to train a classification model. Evaluated on our manually annotated data, our best models reached F1 scores of 65.2 and 59.7 for this task at the levels of elementary and middle school, respectively. We find that discourse level factors contribute to the challenging task of predicting sentence deletion for simplification.