Extraction of CYP Chemical Interactions from Biomedical Literature Using Natural Language Processing Methods

Extraction of CYP Chemical Interactions from Biomedical Literature Using Natural Language Processing Methods
复制标题

DOI:
10.1021/ci800332w
复制
发表时间:
2009-02-01
影响因子:
5.6
通讯作者:
Wild, David J.
Wild, David J.
中科院分区:
化学2区
文献类型:
--
作者:
Jiao, Dazhi;Wild, David J.

文献摘要

被引文献

相似文献

利用自然语言处理和文本挖掘的方法,提出了一个从期刊论文摘要中自动抽取CYP蛋白质和化学作用的系统。在我们的系统中,我们使用了基于最大熵的学习方法,使用了对文本的句法、语义和词汇分析的结果。我们首先介绍了我们的系统结构,然后讨论了用于训练我们的基于机器学习的模型的数据集以及在我们的系统中构建组件的方法,如词性标记(POS)、命名实体识别(NER)、依存关系分析和关系提取。最后对系统进行了测试,得到了很好的结果:系统中的词性、依存关系分析和NER组件达到了很高的准确率,准确率在85.9%到98.5%之间,交互抽取组件的准确率和召回率分别为76.0%和82.6%,整个系统的准确率和召回率分别为68.4%和72.2%。
This paper proposes a system that automatically extracts CYP protein and chemical interactions from journal article abstracts, using natural language processing (NLP) and text mining methods. In our system, we employ a maximum entropy based learning method, using results from syntactic, semantic, and lexical analysis of texts. We first present our system architecture and then discuss the data set for training our machine learning based models and the methods in building components in our system, such as part of speech (POS) tagging, Named Entity Recognition (NER), dependency parsing, and relation extraction. An evaluation of the system is conducted at the end, yielding very promising results: The POS, dependency parsing, and NER components in our system have achieved a very high level of accuracy as measured by precision, ranging from 85.9% to 98.5%, and the precision and the recall of the interaction extraction component are 76.0% and 82.6%, and for the overall system are 68.4% and 72.2%, respectively.