An Ontology-Enabled Natural Language Processing Pipeline for Provenance Metadata Extraction from Biomedical Text (Short Paper).

An Ontology-Enabled Natural Language Processing Pipeline for Provenance Metadata Extraction from Biomedical Text (Short Paper).
复制标题

用于从生物医学文本中提取来源元数据的本体支持的自然语言处理管道(短论文)。

DOI:
10.1007/978-3-319-48472-3_43
复制
发表时间:
2016
期刊:
On the move to meaningful Internet systems ... : CoopIS, DOA, and ODBASE : Confederated International Conferences, CoopIS, DOA, and ODBASE ... proceedings. OTM Confederated International Conferences
影响因子:
--
通讯作者:
Sahoo,SatyaS
Sahoo,SatyaS
中科院分区:
--
文献类型:
--
作者:
Valdez,Joshua;Rueschman,Michael;Kim,Matthew;Redline,Susan;Sahoo,SatyaS

文献摘要

相似文献

由于生物医学领域的复杂性和缺乏适当的自然语言处理(NLP)技术,从生物医学文献中提取结构化信息是一个复杂且具有挑战性的问题。高质量的领域本体以精细的粒度对数据和元数据信息进行建模,可以有效地用于从生物医学文本中准确地提取结构化信息。从已发表的文章中提取描述信息的历史或来源的出处元数据是支持科学再现性的一项重要任务。先前研究报告结果的可重复性是科学进步的基本组成部分。美国国立卫生研究院最近提出的名为“严谨性和可重复性原则”的倡议凸显了这一点。在本文中,我们描述了一种有效的方法,使用支持本体的 NLP 平台从已发表的生物医学研究文献中提取来源元数据,作为临床和医疗保健研究来源 (ProvCaRe) 的一部分。 ProvCaRe-NLP 工具使用起源和生物医学领域本体扩展了临床文本分析和知识提取系统 (cTAKES) 平台。我们使用 20 个同行评审出版物的语料库证明了 ProvCaRe-NLP 工具的有效性。我们的评估结果表明,与 MetaMap 等现有 NLP 流程相比,ProvCaRe-NLP 工具在提取来源元数据方面具有显着更高的召回率。
Extraction of structured information from biomedical literature is a complex and challenging problem due to the complexity of biomedical domain and lack of appropriate natural language processing (NLP) techniques. High quality domain ontologies model both data and metadata information at a fine level of granularity, which can be effectively used to accurately extract structured information from biomedical text. Extraction of provenance metadata, which describes the history or source of information, from published articles is an important task to support scientific reproducibility. Reproducibility of results reported by previous research studies is a foundational component of scientific advancement. This is highlighted by the recent initiative by the US National Institutes of Health called “Principles of Rigor and Reproducibility”. In this paper, we describe an effective approach to extract provenance metadata from published biomedical research literature using an ontology-enabled NLP platform as part of the Provenance for Clinical and Healthcare Research (ProvCaRe). The ProvCaRe-NLP tool extends the clinical Text Analysis and Knowledge Extraction System (cTAKES) platform using both provenance and biomedical domain ontologies. We demonstrate the effectiveness of ProvCaRe-NLP tool using a corpus of 20 peer-reviewed publications. The results of our evaluation demonstrate that the ProvCaRe-NLP tool has significantly higher recall in extracting provenance metadata as compared to existing NLP pipelines such as MetaMap.