Distantly supervised Web relation extraction for knowledge base population

Distantly supervised Web relation extraction for knowledge base population
复制标题

DOI:
10.3233/sw-150180
复制
发表时间:
2016-01-01
期刊:
影响因子:
3
通讯作者:
Ciravegna, Fabio
Ciravegna, Fabio
中科院分区:
计算机科学3区
文献类型:
--
作者:
Augenstein, Isabelle;Maynard, Diana;Ciravegna, Fabio

文献摘要

被引文献

相似文献

从Web页面中提取信息以填充大型跨领域知识库,需要适合跨领域的方法,不需要人工适应新领域,能够处理噪声,并集成从不同Web页面中提取的信息。最近的方法是使用现有的知识库来学习提取信息,并取得了有希望的结果,其中一种方法是远程监督。远程监督是一种无监督的方法,它使用链接开放数据云的背景信息自动标记带有关系的句子,为关系分类器创建训练数据。在本文中,我们提出使用远程监督从Web中提取关系。尽管该方法很有前途,但现有的方法仍然不适合Web提取,因为它们存在三个主要问题:数据稀疏性、噪声和词汇歧义。我们的方法通过使实体识别工具跨域更健壮,并使用无监督的共同引用解析方法跨句子边界提取关系,从而降低了数据稀疏性的影响。采用统计方法对训练数据进行策略性选择,降低了词汇歧义带来的噪声。为了结合从多个来源提取的信息来填充知识库,我们提出并评估了几种信息集成策略,并表明这些策略从使用共同参考分辨率提取的额外关系提及中受益匪浅,精度提高了8%。我们进一步表明,策略性地选择训练数据可以进一步提高3%的精度。
Extracting information from Web pages for populating large, cross-domain knowledge bases requires methods which are suitable across domains, do not require manual effort to adapt to new domains, are able to deal with noise, and integrate information extracted from different Web pages. Recent approaches have used existing knowledge bases to learn to extract information with promising results, one of those approaches being distant supervision. Distant supervision is an unsupervised method which uses background information from the Linking Open Data cloud to automatically label sentences with relations to create training data for relation classifiers. In this paper we propose the use of distant supervision for relation extraction from the Web. Although the method is promising, existing approaches are still not suitable for Web extraction as they suffer from three main issues: data sparsity, noise and lexical ambiguity. Our approach reduces the impact of data sparsity by making entity recognition tools more robust across domains and extracting relations across sentence boundaries using unsupervised co-reference resolution methods. We reduce the noise caused by lexical ambiguity by employing statistical methods to strategically select training data. To combine information extracted from multiple sources for populating knowledge bases we present and evaluate several information integration strategies and show that those benefit immensely from additional relation mentions extracted using co-reference resolution, increasing precision by 8%. We further show that strategically selecting training data can increase precision by a further 3%.