Semi-Supervised Event Extraction with Paraphrase Clusters

Semi-Supervised Event Extraction with Paraphrase Clusters
复制标题

DOI:
10.18653/v1/n18-2058
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
James Ferguson;Colin Lockard;Daniel S. Weld;Hannaneh Hajishirzi
James Ferguson;Colin Lockard;Daniel S. Weld;Hannaneh Hajishirzi
中科院分区:
其他
文献类型:
--
作者:
James Ferguson;Colin Lockard;Daniel S. Weld;Hannaneh Hajishirzi

文献摘要

相似文献

由于缺乏可用的训练数据,监督事件提取系统的准确性受到限制。我们提出了一种自训练事件提取系统的方法,通过引导额外的训练数据。这是通过利用来自多个来源的新闻通讯文章中多次提到同一事件实例来实现的。如果我们的系统可以在这样的集群中对一些提及进行高置信度的提取,那么它就可以通过添加其他提及来获得不同的训练示例。我们的实验表明,在ACE 2005和TAC-KBP 2015数据集上,多个事件提取器的性能得到了显着提高。
Supervised event extraction systems are limited in their accuracy due to the lack of available training data. We present a method for self-training event extraction systems by bootstrapping additional training data. This is done by taking advantage of the occurrence of multiple mentions of the same event instances across newswire articles from multiple sources. If our system can make a high-confidence extraction of some mentions in such a cluster, it can then acquire diverse training examples by adding the other mentions as well. Our experiments show significant performance improvements on multiple event extractors over ACE 2005 and TAC-KBP 2015 datasets.