Extracting Phenotypic Information from the Literature via Natural Language Processing

Extracting Phenotypic Information from the Literature via Natural Language Processing
复制标题

DOI:
10.3233/978-1-60750-949-3-758
复制
发表时间:
2004
影响因子:
--
通讯作者:
Lifeng Chen;C. Friedman
Lifeng Chen;C. Friedman
中科院分区:
--
文献类型:
--
作者:
Lifeng Chen;C. Friedman

文献摘要

相似文献

近年来,生物医学知识的数量呈指数级增长。已经开发了几个自然语言处理(NLP)系统来帮助研究人员从文本文献或叙事报告中自动提取、编码和组织新信息。其中一些系统专注于提取生物实体或分子相互作用,而另一些系统则检索和编码临床信息。为了开发后基因组时代的基因功能,还需要从文献中自动提取表型信息。然而,很少有NLP项目关注这一点。我们介绍了一种名为BioMedLEE的系统的开发,该系统从生物医学文献中提取各种表型信息。该系统是采用现有的临床信息抽取自然语言处理引擎MedLEE开发的。使用随机选择的300个期刊标题进行了一项BioMedLEE的可行性评估研究。结果显示,专家的平均准确率为65.4%(95%CI:[58.0%,72.8%]),召回率为73.0%(95%CI:[66.2%,80.0%])。根据专家协议,BioMedLEE的准确率为64.0%,召回率为77.1%。
In recent years, the amount of biomedical knowledge has been increasing exponentially. Several Natural Language Processing (NLP) systems have been developed to help researchers extract, encode and organize new information automatically from textual literature or narrative reports. Some of these systems focus on extracting biological entities or molecular interactions while others retrieve and encode clinical information. To exploit gene functions in the post-genome era, it is necessary to extract phenotypic information automatically from the literature as well. However, few NLP projects have focused on this. We present the development of a system called BioMedLEE that extracts a broad variety of phenotypic information from the biomedical literature. The system was developed by adapting MedLEE, an existing clinical information extraction NLP engine. A feasibility evaluation study of BioMedLEE was performed using 300 randomly chosen journal titles. Results showed that experts achieved an average precision rate of 65.4%, (95%CI: [58.0%, 72.8%]) and a recall rate of 73.0%, (95%CI: [66.2%, 80.0%]). BioMedLEE had 64.0% precision and 77.1% recall respectively, according to expert agreements.