Protein-protein interaction extraction by leveraging multiple kernels and parsers

Protein-protein interaction extraction by leveraging multiple kernels and parsers
复制标题

DOI:
10.1016/j.ijmedinf.2009.04.010
复制
发表时间:
2009-12-01
影响因子:
4.9
通讯作者:
Tsujii, Jun'ichi
Tsujii, Jun'ichi
中科院分区:
医学2区
文献类型:
--
作者:
Miwa, Makoto;Saetre, Rune;Tsujii, Jun'ichi

文献摘要

被引文献

相似文献

蛋白质间相互作用抽取是生物医学自然语言处理领域的一个重要研究课题。基于核的机器学习方法已经被广泛用于自动提取PPI,并且已经发布了几种专注于句子结构不同部分的核用于PPI任务。在本文中,我们提出了一种方法,联合收割机内核的基础上,几个句法分析器,以检索最广泛的重要信息,从一个给定的句子。我们使用支持向量机(SVM)的方法进行评估,我们取得了更好的结果比其他国家的最先进的PPI系统上的五个语料库。此外,我们从PPI提取的角度分析了五个语料库的兼容性,我们看到其中一些语料库有很小的不兼容性,但他们仍然可以通过一点努力来组合。(C)2009爱思唯尔爱尔兰有限公司保留所有权利。
Protein-protein interaction (PPI) extraction is an important and widely researched task in the biomedical natural language processing (BioNLP) field. Kernel-based machine learning methods have been used widely to extract PPI automatically, and several kernels focusing on different parts of sentence structure have been published for the PPI task. In this paper, we propose a method to combine kernels based on several syntactic parsers, in order to retrieve the widest possible range of important information from a given sentence. We evaluate the method using a support vector machine (SVM), and we achieve better results than other state-of-the-art PPI systems on four out of five corpora. Further, we analyze the compatibility of the five corpora from the viewpoint of PPI extraction, and we see that some of them have small incompatibilities, but they can still be combined with a little effort. (C) 2009 Elsevier Ireland Ltd. All rights reserved.