A Hybrid Framework for Semantic Relation Extraction over Enterprise Data

A Hybrid Framework for Semantic Relation Extraction over Enterprise Data
复制标题

企业数据语义关系提取的混合框架

DOI:
10.4018/ijswis.2015070101
复制
发表时间:
2015
影响因子:
3.2
通讯作者:
Wang Min
Wang Min
中科院分区:
计算机科学4区
文献类型:
--
作者:
Shen Wei;Wang Jianyong;Luo Ping;Wang Min

文献摘要

相似文献

近年来,Web数据中的关系抽取引起了人们的广泛关注。然而,当涉及到从企业数据中提取关系时,几乎没有做任何工作,而不管在真实的应用中对这种工作的迫切需要,电子取证与Web数据相比,企业数据的一个显著特点是其冗余度低。以往从Web数据中提取关系的工作主要依赖于数据的高冗余度,因此不能有效地应用于企业数据。本文提出了一个无监督的混合框架,称为反应器。REACTOR结合了统计方法、分类和聚类,自动识别企业数据中出现的实体之间的各种类型的关系。此外,作者还尝试使用代词回指消解来提取更多的跨句子关系。他们在来自HP的包含超过300万页的真实企业数据集上评估了REACTOR,实验结果显示了REACTOR的有效性。
Relation extraction from the Web data has attracted a lot of attention in recent years. However, little work has been done when it comes to relation extraction from the enterprise data regardless of the urgent needs to such work in real applications e.g., E-discovery. One distinct characteristic of the enterprise data in comparison with the Web data is its low redundancy. Previous work on relation extraction from the Web data largely relies on the data's high redundancy level and thus cannot be applied to the enterprise data effectively. This paper proposes an unsupervised hybrid framework called REACTOR. REACTOR combines a statistical method, classification, and clustering to identify various types of relations among entities appearing in the enterprise data automatically. Furthermore, the authors explore to apply pronominal anaphora resolution to extract more relations expressed across multiple sentences. They evaluate REACTOR over a real-world enterprise data set from HP that contains over three million pages and the experimental results show the effectiveness of REACTOR.