Applying the Naive Bayes classifier with kernel density estimation to the prediction of protein-protein interaction sites

Applying the Naive Bayes classifier with kernel density estimation to the prediction of protein-protein interaction sites
复制标题

DOI:
10.1093/bioinformatics/btq302
复制
发表时间:
2010-08-01
期刊:
影响因子:
5.8
通讯作者:
Mizuguchi, Kenji
Mizuguchi, Kenji
中科院分区:
生物学3区
文献类型:
--
作者:
Murakami, Yoichi;Mizuguchi, Kenji

文献摘要

被引文献

相似文献

动机:蛋白质结构的有限性往往限制了蛋白质的功能注释和蛋白质相互作用位点的鉴定。因此,需要通过计算方法从蛋白质序列中识别相互作用位点,以揭示许多蛋白质的功能。本文介绍了一种预测蛋白质序列中相互作用位点的新方法(PSIVER)。仅序列特征(位置特异性评分矩阵和预测可达性)用于训练朴素贝叶斯分类器(NBC),并且使用核密度估计方法(KDE)估计每个序列特征的条件概率。PSIVER的留一交叉验证实现了0.151的马修斯相关系数(MCC),35.3%的F-测量,从蛋白质数据库中的105个异源二聚体中提取的186个蛋白质序列的非冗余集(由36219个残基组成,其中15.2%是已知的界面残基)上的精确度为30.6%,召回率为41.6%。尽管用于训练的数据集是高度不平衡的,但随机化测试表明,所提出的方法能够避免过拟合。PSIVER还在72个未用于训练的序列(由18140个残基组成,其中10.6%是已知的界面残基)上进行了测试,并实现了0.135的MCC,31.5%的F-测量,25.0%的精确度和46.5%的召回率,优于在相同数据集上测试的其他公开可用的服务器。PSIVER使实验生物学家能够仅从序列信息中识别未知蛋白质中的潜在界面残基,并选择性地突变这些残基以解开蛋白质功能。
Motivation: The limited availability of protein structures often restricts the functional annotation of proteins and the identification of their protein-protein interaction sites. Computational methods to identify interaction sites from protein sequences alone are, therefore, required for unraveling the functions of many proteins. This article describes a new method (PSIVER) to predict interaction sites, i.e. residues binding to other proteins, in protein sequences. Only sequence features (position-specific scoring matrix and predicted accessibility) are used for training a Naive Bayes classifier (NBC), and conditional probabilities of each sequence feature are estimated using a kernel density estimation method (KDE).Results: The leave-one out cross-validation of PSIVER achieved a Matthews correlation coefficient (MCC) of 0.151, an F-measure of 35.3%, a precision of 30.6% and a recall of 41.6% on a non-redundant set of 186 protein sequences extracted from 105 heterodimers in the Protein Data Bank (consisting of 36 219 residues, of which 15.2% were known interface residues). Even though the dataset used for training was highly imbalanced, a randomization test demonstrated that the proposed method managed to avoid overfitting. PSIVER was also tested on 72 sequences not used in training (consisting of 18 140 residues, of which 10.6% were known interface residues), and achieved an MCC of 0.135, an F-measure of 31.5%, a precision of 25.0% and a recall of 46.5%, outperforming other publicly available servers tested on the same dataset. PSIVER enables experimental biologists to identify potential interface residues in unknown proteins from sequence information alone, and to mutate those residues selectively in order to unravel protein functions.