Exploiting semantic annotations for open information extraction: an experience in the biomedical domain

Exploiting semantic annotations for open information extraction: an experience in the biomedical domain
复制标题

DOI:
10.1007/s10115-012-0590-x
复制
发表时间:
2014-02-01
影响因子:
2.7
通讯作者:
Berlanga, Rafael
Berlanga, Rafael
中科院分区:
计算机科学4区
文献类型:
--
作者:
Nebot, Victoria;Berlanga, Rafael

文献摘要

被引文献

相似文献

Web上发布的非结构化文本的数量越来越多,这就需要新的工具和方法来自动处理和提取相关信息。传统的信息提取集中于获取特定于领域的预先指定的关系,这通常需要体力劳动和重型机械;特别是在生物医学领域,主要工作是识别定义良好的实体,如基因或蛋白质,这构成了提取识别实体之间关系的基础。Web的内在特征和规模需要新的方法来处理文档的多样性,其中关系的数量是无限的,并且事先不知道。本文提出了一种利用语义标注中的知识从文本中提取与领域无关的关系的可扩展方法。该方法不针对任何特定领域(例如,蛋白质-蛋白质相互作用和药物-药物相互作用),不需要任何人工输入或深度处理。此外,该方法使用提取的关系来计算以签名类型和同义关系串为特征的抽象语义关系组。在构建正式知识库时,这构成了一个有价值的知识来源,因为我们能够通过语义标注过程将提取的关系与可用的知识资源无缝集成。所提出的方法已经成功地应用于生物医学领域的大量文本集合,结果非常令人鼓舞。
The increasing amount of unstructured text published on the Web is demanding new tools and methods to automatically process and extract relevant information. Traditional information extraction has focused on harvesting domain-specific, pre-specified relations, which usually requires manual labor and heavy machinery; especially in the biomedical domain, the main efforts have been directed toward the recognition of well-defined entities such as genes or proteins, which constitutes the basis for extracting the relationships between the recognized entities. The intrinsic features and scale of the Web demand new approaches able to cope with the diversity of documents, where the number of relations is unbounded and not known in advance. This paper presents a scalable method for the extraction of domain-independent relations from text that exploits the knowledge in the semantic annotations. The method is not geared to any specific domain (e.g., protein-protein interactions and drug-drug interactions) and does not require any manual input or deep processing. Moreover, the method uses the extracted relations to compute groups of abstract semantic relations characterized by their signature types and synonymous relation strings. This constitutes a valuable source of knowledge when constructing formal knowledge bases, as we enable seamless integration of the extracted relations with the available knowledge resources through the process of semantic annotation. The proposed approach has successfully been applied to a large text collection in the biomedical domain and the results are very encouraging.