The interaction of domain knowledge and linguistic structure in natural language processing: interpreting hypernymic propositions in biomedical text

The interaction of domain knowledge and linguistic structure in natural language processing: interpreting hypernymic propositions in biomedical text
复制标题

DOI:
10.1016/j.jbi.2003.11.003
复制
发表时间:
2003-12-01
影响因子:
4.5
通讯作者:
Fiszman, M
Fiszman, M
中科院分区:
医学3区
文献类型:
--
作者:
Rindflesch, TC;Fiszman, M

文献摘要

被引文献

相似文献

对自由文本文档(如MEDLINE引文)中的语义命题的解释将为生物医学应用提供有价值的支持,生物医学信息学社区正在研究几种语义解释的方法。在本文中,我们描述了一种解释语言结构编码的方法,其中一个更具体的概念与一个更一般的概念在分类关系中。为了有效地处理这些结构,我们利用了来自统一医学语言系统(UMLS)的未指定语法分析和结构化领域知识。在介绍了我们的系统所依赖的语法处理之后,我们将重点关注支持高义命题解释的UMLS知识。我们首先使用语义网络中的语义组来确保涉及的两个概念是兼容的;然后,元词典中的层次信息决定了哪个概念更一般,哪个更具体。基于语义组Chemicals and Drugs对样本的初步评估提供了83%的精度。进行了误差分析,并提出了可能的解决方法。本文的研究为研究自然语言处理中领域知识与语言结构之间的相互作用提供了一个范例,也为篇章结构的自动处理研究做出了贡献。我们提出的系统的其他含义包括其集成在生物医学文本的高级语义解释处理器中,并用于特定领域的信息提取。该方法具有支持一系列应用的潜力,包括信息检索和本体工程。Elsevier Inc.出版。
Interpretation of semantic propositions in free-text documents such as MEDLINE citations would provide valuable support for biomedical applications, and several approaches to semantic interpretation are being pursued in the biomedical informatics community. In this paper, we describe a methodology for interpreting linguistic structures that encode hypernymic propositions, in which a more specific concept is in a taxonomic relationship with a more general concept. In order to effectively process these constructions, we exploit underspecified syntactic analysis and structured domain knowledge from the Unified Medical Language System (UMLS). After introducing the syntactic processing on which our system depends, we focus on the UMLS knowledge that supports interpretation of hypernymic propositions. We first use semantic groups from the Semantic Network to ensure that the two concepts involved are compatible; hierarchical information in the Metathesaurus then determines which concept is more general and which more specific. A preliminary evaluation of a sample based on the semantic group Chemicals and Drugs provides 83% precision. An error analysis was conducted and potential solutions to the problems encountered are presented. The research discussed here serves as a paradigm for investigating the interaction between domain knowledge and linguistic structure in natural language processing, and could also make a contribution to research on automatic processing of discourse structure. Additional implications of the system we present include its integration in advanced semantic interpretation processors for biomedical text and its use for information extraction in specific domains. The approach has the potential to support a range of applications, including information retrieval and ontology engineering. Published by Elsevier Inc.