Leveraging Linguistic Structure For Open Domain Information Extraction

Leveraging Linguistic Structure For Open Domain Information Extraction
复制标题

DOI:
10.3115/v1/p15-1034
复制
发表时间:
2015-07
期刊:
--
影响因子:
--
通讯作者:
Gabor Angeli;Melvin Johnson;Christopher D. Manning
Gabor Angeli;Melvin Johnson;Christopher D. Manning
中科院分区:
其他
文献类型:
--
作者:
Gabor Angeli;Melvin Johnson;Christopher D. Manning

文献摘要

被引文献

相似文献

由开放域信息抽取(Open Domain Information Extraction,Open IE)系统产生的关系三元组对于问答、推理和其他IE任务是有用的。传统上,这些都是使用大量的模式来提取的;然而,这种方法在域外文本和长期依赖关系上很脆弱,并且没有深入了解参数的子结构。我们用一些规范结构的句子的模式来替换这个大的模式集,并将焦点转移到一个分类器上,该分类器学习从较长的句子中提取自包含的子句。然后,我们对这些短子句进行自然逻辑推理,以确定每个候选三元组的最大具体参数。我们表明,我们的方法优于一个国家的最先进的开放IE系统的端到端的TAC-KBP 2013插槽填充任务。
Relation triples produced by open domain information extraction (open IE) systems are useful for question answering, inference, and other IE tasks. Traditionally these are extracted using a large set of patterns; however, this approach is brittle on out-of-domain text and long-range dependencies, and gives no insight into the substructure of the arguments. We replace this large pattern set with a few patterns for canonically structured sentences, and shift the focus to a classifier which learns to extract self-contained clauses from longer sentences. We then run natural logic inference over these short clauses to determine the maximally specific arguments for each candidate triple. We show that our approach outperforms a state-of-the-art open IE system on the end-to-end TAC-KBP 2013 Slot Filling task.