Evaluation of different biological data and computational classification methods for use in protein interaction prediction

Evaluation of different biological data and computational classification methods for use in protein interaction prediction
复制标题

DOI:
10.1002/prot.20865
复制
发表时间:
2006-05-15
影响因子:
2.9
通讯作者:
Klein-Seetharaman, J
Klein-Seetharaman, J
中科院分区:
生物学4区
文献类型:
--
作者:
Qi, YJ;Bar-Joseph, Z;Klein-Seetharaman, J

文献摘要

被引文献

相似文献

蛋白质-蛋白质相互作用在许多生物系统中起着关键作用。高通量方法可以直接检测酵母中相互作用的蛋白质组,但结果往往是不完整的,并表现出较高的假阳性和假阴性率。最近,许多不同的研究小组独立地建议使用监督学习方法来整合蛋白质相互作用预测任务的直接和间接生物数据源。然而,数据源、方法和实现方式各不相同。此外,蛋白质相互作用预测任务本身可以细分为(1)物理相互作用、(2)共复合物关系和(3)途径共成员关系的预测。为了系统地研究不同数据源的效用以及将数据编码为预测每种类型蛋白质相互作用的特征的方式,我们收集了大量的生物特征,并改变了它们的编码以用于三个预测任务中的每一个。使用六种不同的分类器来评估预测相互作用的准确性,随机森林(RF),基于RF相似性的k-最近邻,朴素贝叶斯,决策树,逻辑回归和支持向量机。对于所有的分类器,三个预测任务有不同的成功率,和共同复杂的预测似乎是一个更容易的任务比其他两个。然而,独立于预测任务,RF分类器始终被列为所有特征集组合的前两个分类器之一。因此,我们使用该分类器来研究不同生物数据集的重要性。首先,我们使用RF树结构的分裂函数,基尼指数,来估计特征的重要性。其次,我们确定了分类精度时,只有排名靠前的功能被用作分类器中的输入。我们发现,不同特征的重要性取决于具体的预测任务和它们的编码方式。引人注目的是,基因表达始终是所有三个预测任务的最重要特征,而使用酵母双杂交系统确定的蛋白质相互作用在任何条件下都不是最重要的特征。
Protein-protein interactions play a key role in many biological systems. High-throughput methods can directly detect the set of interacting proteins in yeast, but the results are often incomplete and exhibit high false-positive and false-negative rates. Recently, many different research groups independently suggested using supervised learning methods to integrate direct and indirect biological data sources for the protein interaction prediction task. However, the data sources, approaches, and implementations varied. Furthermore, the protein interaction prediction task itself can be subdivided into prediction of (1) physical interaction, (2) co-complex relationship, and (3) pathway co-membership. To investigate systematically the utility of different data sources and the way the data is encoded as features for predicting each of these types of protein interactions, we assembled a large set of biological features and varied their encoding for use in each of the three prediction tasks. Six different classifiers were used to assess the accuracy in predicting interactions, Random Forest (RF), RF similarity-based k-Nearest-Neighbor, Naive Bayes, Decision Tree, Logistic Regression, and Support Vector Machine. For all classifiers, the three prediction tasks had different success rates, and co-complex prediction appears to be an easier task than the other two. Independently of prediction task, however, the RF classifier consistently ranked as one of the top two classifiers for all combinations of feature sets. Therefore, we used this classifier to study the importance of different biological datasets. First, we used the splitting function of the RF tree structure, the Gini index, to estimate feature importance. Second, we determined classification accuracy when only the top-ranking features were used as an input in the classifier. We find that the importance of different features depends on the specific prediction task and the way they are encoded. Strikingly, gene expression is consistently the most important feature for all three prediction tasks, while the protein interactions identified using the yeast-2-hybrid system were not among the top-ranking features under any condition.