Learning to Simplify Sentences with Quasi-Synchronous Grammar and Integer Programming

Learning to Simplify Sentences with Quasi-Synchronous Grammar and Integer Programming
复制标题

DOI:
--
复制
发表时间:
2011-07
期刊:
--
影响因子:
--
通讯作者:
K. Woodsend;Mirella Lapata
K. Woodsend;Mirella Lapata
中科院分区:
其他
文献类型:
--
作者:
K. Woodsend;Mirella Lapata

文献摘要

被引文献

相似文献

文本简化旨在将文本重写为更简单的版本,从而使更广泛的受众能够获取信息。之前的大多数工作都使用旨在分割长句子的手工规则来简化句子,或者使用预定义的字典来替换困难的单词。本文提出了一种基于准同步语法的数据驱动模型,这是一种可以自然捕获结构不匹配和复杂重写操作的形式主义。我们描述了如何从维基百科中归纳出这样的语法,并提出了一种整数线性规划模型,用于从语法生成的可能重写空间中选择最合适的简化。我们通过实验证明,我们的方法进行了简化,显着降低了输入的阅读难度,同时保持了语法性并保留了其含义。
Text simplification aims to rewrite text into simpler versions, and thus make information accessible to a broader audience. Most previous work simplifies sentences using handcrafted rules aimed at splitting long sentences, or substitutes difficult words using a predefined dictionary. This paper presents a data-driven model based on quasi-synchronous grammar, a formalism that can naturally capture structural mismatches and complex rewrite operations. We describe how such a grammar can be induced from Wikipedia and propose an integer linear programming model for selecting the most appropriate simplification from the space of possible rewrites generated by the grammar. We show experimentally that our method creates simplifications that significantly reduce the reading difficulty of the input, while maintaining grammaticality and preserving its meaning.