Seed Selection for Distantly Supervised Web-Based Relation Extraction

Seed Selection for Distantly Supervised Web-Based Relation Extraction
复制标题

DOI:
10.3115/v1/w14-6203
复制
发表时间:
2014-08
期刊:
--
影响因子:
--
通讯作者:
Isabelle Augenstein
Isabelle Augenstein
中科院分区:
其他
文献类型:
--
作者:
Isabelle Augenstein

文献摘要

相似文献

在本文中,我们考虑远程监督问题,通过使用链接开放数据云中的背景信息自动标记 Web 文档,然后将其用作关系分类器的训练数据,从网页中提取某些类别(例如音乐艺术家)的实体(例如“披头士”)的关系(例如起源(音乐艺术家、位置))。远程监督方法通常会遇到自动标记文本时的歧义问题,以及判断提及是否是真实关系提及的背景数据不完整的问题。本文探讨了这样一个假设:基于背景数据的简单统计方法可以帮助过滤不可靠的训练数据,从而提高关系提取器的精度。在网络语料库上的实验表明,通过策略性地选择种子数据可以将错误减少 35%。
In this paper we consider the problem of distant supervision to extract relations (e.g. origin(musical artist, location)) for entities (e.g. ‘The Beatles’) of certain classes (e.g. musical artist) from Web pages by using background information from the Linking Open Data cloud to automatically label Web documents which are then used as training data for relation classifiers. Distant supervision approaches typically su er from the problem of ambiguity when automatically labelling text, as well as the problem of incompleteness of background data to judge whether a mention is a true relation mention. This paper explores the hypothesis that simple statistical methods based on background data can help to filter unreliable training data and thus improve the precision of relation extractors. Experiments on a Web corpus show that an error reduction of 35% can be achieved by strategically selecting seed data.