Text Rewriting Improves Semantic Role Labeling

Text Rewriting Improves Semantic Role Labeling
复制标题

DOI:
10.1613/jair.4431
复制
发表时间:
2014-09
期刊:
J. Artif. Intell. Res.
影响因子:
--
通讯作者:
K. Woodsend;Mirella Lapata
K. Woodsend;Mirella Lapata
中科院分区:
其他
文献类型:
--
作者:
K. Woodsend;Mirella Lapata

文献摘要

相似文献

大规模标注语料库是开发高性能NLP系统的先决条件。这样的语料库制作成本高,规模有限,往往需要语言专业知识。在本文中,我们使用文本重写作为增加可用于模型训练的标记数据量的一种手段。我们的方法使用自动提取的重写规则,从可比的语料库和双文本生成多个版本的句子注释的黄金标准标签。我们将这一想法应用于语义角色标记,并表明在重写数据上训练的模型在CoNLL-2009基准数据集上的表现优于最先进的模型。
Large-scale annotated corpora are a prerequisite to developing high-performance NLP systems. Such corpora are expensive to produce, limited in size, often demanding linguistic expertise. In this paper we use text rewriting as a means of increasing the amount of labeled data available for model training. Our method uses automatically extracted rewrite rules from comparable corpora and bitexts to generate multiple versions of sentences annotated with gold standard labels. We apply this idea to semantic role labeling and show that a model trained on rewritten data outperforms the state of the art on the CoNLL-2009 benchmark dataset.